Preview

Arctic XXI century

Advanced search

A model of a universal linguistic knowledge graph for Turkic languages

https://doi.org/10.25587/3034-7378-2026-14-3-5-21

Abstract

The article describes a new model of a universal linguistic knowledge graph for Turkic languages as a basis for creating ontolinguistic resources. In the model, world knowledge is formalized as an ontological core that integrates taxonomic and situational ontologies (by analogy with WordNet and FrameNet) and is closely interrelated with linguistic data from Turkic languages. The relevance of the work is determined by the low-resource nature of most Turkic languages and the loss of effectiveness of neural network approaches in processing agglutinative languages with rich morphology. The goal of the project is to develop a new model of a universal linguistic knowledge graph for Turkic languages. The objectives are: integrating world knowledge and deep linguistic information at all levels of the language system; constructing the graph according to the principle of a semiotic vertical; ensuring the morphemocentricity of the model; and integrating it with international language standards. The foundational method is a pragmatically-oriented approach that focuses the research domain on the structural and functional features of Turkic languages. The knowledge graph is constructed according to the principle of a semiotic vertical that unites five levels – phonological, morphological, syntactic, semantic, and pragmatic – as well as three inter-level bridges (morphonological, morphosyntactic, and morphosemantic). A key feature of the model is its morphemocentricity: the morpheme acts as the basic linguistic unit, connected with units of all language levels of the model, which fundamentally distinguishes this approach from word-centric models for fusional languages. The model is integrated with international standards: IPA, Universal Dependencies, UDS, UMR, FrameNet, WordNet, and DiAML, each of which is used to describe linguistic elements and relations at its own level. The model has already found practical application in several software products, such as the linguistic portal “Turkic Morpheme” and the new version of the Tatar corpus “Tugan Tel”. Prospects include improving the interpretability of LLMs, improving the quality of machine translation, and unifying annotations of electronic corpora. 

About the Authors

A. R. Gatiatullin
Institute of Applied Semiotics of Tatarstan Academy of Sciences
Russian Federation

Ayrat R. GATIATULLIN - Cand. Sci. (Engineering), Leading Researcher

Kazan

WoS ResearcherID: ABF8777-2020, Scopus Author ID: 56500678000, Elibrary AuthorID: 161758



N. A. Prokopyev
Institute of Applied Semiotics of Tatarstan Academy of Sciences
Russian Federation

Nikolai A. PROKOPYEV - Cand. Sci. (Engineering), Researcher

Kazan

WoS ResearcherID: S-3829-2016, Scopus Author ID: 57206891311, Elibrary AuthorID: 999214



N. Z. Abdurakhmonova
National University of Uzbekistan Named after Mirzo Ulugbek
Uzbekistan

Nilufar Z. ABDURAKHMONOVA - Dr. Sci. (Philology), Professor, Head of the Department of Computational and Applied Linguistics, Faculty of Journalism and Uzbek Philology

Tashkent

WoS ResearcherID: ABC-5275-2021, Scopus Author ID: 57221106516



A. А. Kasieva
Kyrgyz-Turkish Manas University
Kyrgyzstan

Aida A. KASIEVA - Cand. Sci. (Philology), Associate Professor, Professor and Head of the Department of Translations, Faculty of Humanities

Bishkek

WoS ResearcherID: HGC-3155-2022, Scopus Author ID: 57189928164



References

1. Lewis P, Perez E, Piktus A et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. NIPS’20: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020:9459–9474. DOI: https://doi.org/10.5555/3495724.3496517

2. Edge D, Trinh H, Cheng N et al. From local to global: a graph RAG approach to query-focused summarization. arXiv cs.CL 2404.16130. 2025. DOI: https://doi.org/10.48550/arXiv.2404.16130

3. Lomov PA. Using ontologies to contextualize queries to Large Language Models. Ontology of Designing. 2025;15(2):239–248. DOI: https://doi.org/10.18287/2223-9537-2025-15-2-239–248 (in Russian).

4. Arnett C, Bergen B. Why do language models perform worse for morphologically complex languages? Proceedings of the 31st International Conference on Computational Linguistics. UAE. 2025:6607–6623.

5. Yu S, Kulkarni N, Lee H, and Kim J. Syllable-level Neural Language Model for agglutinative language. Proceedings of the First Workshop on Subword and Character Level Models in NLP. Copenhagen. 2017:92–96.

6. Guzev VG. About some exotic features of Turkic languages (“Turkic miracles”). Digest of World Politics. 2020;(10):231–245 (in Russian).

7. McCrae J, Spohr D, Cimiano P. Linking lexical resources and ontologies on the Semantic Web with Lemon. The Semantic Web: Research and Applications. ESWC 2011. Lecture Notes in Computer Science. Berlin, Heidelberg: Springer. 2011;(6643):245–259. DOI: https://doi.org/10.1007/978-3-642-21034-1_17

8. Francopoulo G, George M, Calzolari N, Monachini M, Bel N, Pet M, and Soria C. Lexical Markup Framework (LMF). Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06). Italy. 2006.

9. Gatiatullin AR, Prokopyev NA, Suleymanov DSh. Model of linguistic knowledge graphs of Turkic languages. Ontologiya proyektirovaniya. 2024;14(3):366–378 (in Russian).

10. Linden K, Axelson E, Drobac S et al. HFST – A system for creating NLP tools. Systems and Frameworks for Computational Morphology. SFCM 2013. Communications in Computer and Information Science. 2013;(380):53–71. DOI: https://doi.org/10.1007/978-3-642-40486-3_4

11. Marneffe M, Manning C, Nivre J et al. Universal Dependencies. Computational Linguistics. 2021;47(2):255–308. DOI: https://doi.org/10.1162/coli_a_00402

12. Fellbaum C. WordNet: an electronic lexical database. Cambridge. MA: MIT Press. 1998. DOI: https://doi.org/10.7551/mitpress/7287.001.0001

13. Baker CF, Fillmore CJ, Lowe JB. The Berkeley FrameNet project. 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics. 1998;(1):86–90. DOI: https://doi.org/10.3115/980845.980860

14. Van Gysel JEL, Vigus M, Chun J et al. Designing a Uniform Meaning Representation for Natural Language Processing. Künstliche Intelligenz. 2021;(35):343–360. DOI: https://doi.org/10.1007/s13218-021-00722-w

15. Banarescu L, Bonial C, Cai S et al. Abstract Meaning Representation for Sembanking. Proceedings of the 7th Linguistic Annotation Workshop & Interoperability with Discourse. 2013:178–186.

16. Bunt H, Alexandersson J, Choe JW et al. ISO 24617-2: A semantically-based standard for dialogue annotation. Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12). 2012:430–437.

17. Gatiatullin A, Suleymanov D, Prokopyev N, Khakimov B. About “Turkic Morpheme” portal. CEUR Workshop Proceedings. 2020;(2780):226–243.

18. Mukhamedshin D, Gatiatullin A, Gilmullin R. The new version of the corpus data management system “Tugan Tel” using graph knowledge base. IEEE 3rd International Conference on Problems of Informatics, Electronics and Radio Engineering (PIERE). 2024:1800–1804. DOI: https://doi.org/10.1109/PIERE62470.2024.10804932

19. White AS, Reisinger D, Sakagucih K et al. Universal Decompositional Semantics on Universal Dependencies. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016:1713–1723. DOI: https://doi.org/10.18653/v1/D16-1177


Review

For citations:


Gatiatullin A.R., Prokopyev N.A., Abdurakhmonova N.Z., Kasieva A.А. A model of a universal linguistic knowledge graph for Turkic languages. Arctic XXI century. 2026;(3):5-21. (In Russ.) https://doi.org/10.25587/3034-7378-2026-14-3-5-21

Views: 18

JATS XML

ISSN 3034-7378 (Print)
ISSN 3034-7386 (Online)