A model of a universal linguistic knowledge graph for Turkic languages
https://doi.org/10.25587/3034-7378-2026-14-3-5-21
Abstract
The article describes a new model of a universal linguistic knowledge graph for Turkic languages as a basis for creating ontolinguistic resources. In the model, world knowledge is formalized as an ontological core that integrates taxonomic and situational ontologies (by analogy with WordNet and FrameNet) and is closely interrelated with linguistic data from Turkic languages. The relevance of the work is determined by the low-resource nature of most Turkic languages and the loss of effectiveness of neural network approaches in processing agglutinative languages with rich morphology. The goal of the project is to develop a new model of a universal linguistic knowledge graph for Turkic languages. The objectives are: integrating world knowledge and deep linguistic information at all levels of the language system; constructing the graph according to the principle of a semiotic vertical; ensuring the morphemocentricity of the model; and integrating it with international language standards. The foundational method is a pragmatically-oriented approach that focuses the research domain on the structural and functional features of Turkic languages. The knowledge graph is constructed according to the principle of a semiotic vertical that unites five levels – phonological, morphological, syntactic, semantic, and pragmatic – as well as three inter-level bridges (morphonological, morphosyntactic, and morphosemantic). A key feature of the model is its morphemocentricity: the morpheme acts as the basic linguistic unit, connected with units of all language levels of the model, which fundamentally distinguishes this approach from word-centric models for fusional languages. The model is integrated with international standards: IPA, Universal Dependencies, UDS, UMR, FrameNet, WordNet, and DiAML, each of which is used to describe linguistic elements and relations at its own level. The model has already found practical application in several software products, such as the linguistic portal “Turkic Morpheme” and the new version of the Tatar corpus “Tugan Tel”. Prospects include improving the interpretability of LLMs, improving the quality of machine translation, and unifying annotations of electronic corpora.
About the Authors
A. R. GatiatullinRussian Federation
Ayrat R. GATIATULLIN - Cand. Sci. (Engineering), Leading Researcher
Kazan
WoS ResearcherID: ABF8777-2020, Scopus Author ID: 56500678000, Elibrary AuthorID: 161758
N. A. Prokopyev
Russian Federation
Nikolai A. PROKOPYEV - Cand. Sci. (Engineering), Researcher
Kazan
WoS ResearcherID: S-3829-2016, Scopus Author ID: 57206891311, Elibrary AuthorID: 999214
N. Z. Abdurakhmonova
Uzbekistan
Nilufar Z. ABDURAKHMONOVA - Dr. Sci. (Philology), Professor, Head of the Department of Computational and Applied Linguistics, Faculty of Journalism and Uzbek Philology
Tashkent
WoS ResearcherID: ABC-5275-2021, Scopus Author ID: 57221106516
A. А. Kasieva
Kyrgyzstan
Aida A. KASIEVA - Cand. Sci. (Philology), Associate Professor, Professor and Head of the Department of Translations, Faculty of Humanities
Bishkek
WoS ResearcherID: HGC-3155-2022, Scopus Author ID: 57189928164
References
1. Lewis P, Perez E, Piktus A et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. NIPS’20: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020:9459–9474. DOI: https://doi.org/10.5555/3495724.3496517
2. Edge D, Trinh H, Cheng N et al. From local to global: a graph RAG approach to query-focused summarization. arXiv cs.CL 2404.16130. 2025. DOI: https://doi.org/10.48550/arXiv.2404.16130
3. Lomov PA. Using ontologies to contextualize queries to Large Language Models. Ontology of Designing. 2025;15(2):239–248. DOI: https://doi.org/10.18287/2223-9537-2025-15-2-239–248 (in Russian).
4. Arnett C, Bergen B. Why do language models perform worse for morphologically complex languages? Proceedings of the 31st International Conference on Computational Linguistics. UAE. 2025:6607–6623.
5. Yu S, Kulkarni N, Lee H, and Kim J. Syllable-level Neural Language Model for agglutinative language. Proceedings of the First Workshop on Subword and Character Level Models in NLP. Copenhagen. 2017:92–96.
6. Guzev VG. About some exotic features of Turkic languages (“Turkic miracles”). Digest of World Politics. 2020;(10):231–245 (in Russian).
7. McCrae J, Spohr D, Cimiano P. Linking lexical resources and ontologies on the Semantic Web with Lemon. The Semantic Web: Research and Applications. ESWC 2011. Lecture Notes in Computer Science. Berlin, Heidelberg: Springer. 2011;(6643):245–259. DOI: https://doi.org/10.1007/978-3-642-21034-1_17
8. Francopoulo G, George M, Calzolari N, Monachini M, Bel N, Pet M, and Soria C. Lexical Markup Framework (LMF). Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06). Italy. 2006.
9. Gatiatullin AR, Prokopyev NA, Suleymanov DSh. Model of linguistic knowledge graphs of Turkic languages. Ontologiya proyektirovaniya. 2024;14(3):366–378 (in Russian).
10. Linden K, Axelson E, Drobac S et al. HFST – A system for creating NLP tools. Systems and Frameworks for Computational Morphology. SFCM 2013. Communications in Computer and Information Science. 2013;(380):53–71. DOI: https://doi.org/10.1007/978-3-642-40486-3_4
11. Marneffe M, Manning C, Nivre J et al. Universal Dependencies. Computational Linguistics. 2021;47(2):255–308. DOI: https://doi.org/10.1162/coli_a_00402
12. Fellbaum C. WordNet: an electronic lexical database. Cambridge. MA: MIT Press. 1998. DOI: https://doi.org/10.7551/mitpress/7287.001.0001
13. Baker CF, Fillmore CJ, Lowe JB. The Berkeley FrameNet project. 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics. 1998;(1):86–90. DOI: https://doi.org/10.3115/980845.980860
14. Van Gysel JEL, Vigus M, Chun J et al. Designing a Uniform Meaning Representation for Natural Language Processing. Künstliche Intelligenz. 2021;(35):343–360. DOI: https://doi.org/10.1007/s13218-021-00722-w
15. Banarescu L, Bonial C, Cai S et al. Abstract Meaning Representation for Sembanking. Proceedings of the 7th Linguistic Annotation Workshop & Interoperability with Discourse. 2013:178–186.
16. Bunt H, Alexandersson J, Choe JW et al. ISO 24617-2: A semantically-based standard for dialogue annotation. Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12). 2012:430–437.
17. Gatiatullin A, Suleymanov D, Prokopyev N, Khakimov B. About “Turkic Morpheme” portal. CEUR Workshop Proceedings. 2020;(2780):226–243.
18. Mukhamedshin D, Gatiatullin A, Gilmullin R. The new version of the corpus data management system “Tugan Tel” using graph knowledge base. IEEE 3rd International Conference on Problems of Informatics, Electronics and Radio Engineering (PIERE). 2024:1800–1804. DOI: https://doi.org/10.1109/PIERE62470.2024.10804932
19. White AS, Reisinger D, Sakagucih K et al. Universal Decompositional Semantics on Universal Dependencies. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016:1713–1723. DOI: https://doi.org/10.18653/v1/D16-1177
Review
For citations:
Gatiatullin A.R., Prokopyev N.A., Abdurakhmonova N.Z., Kasieva A.А. A model of a universal linguistic knowledge graph for Turkic languages. Arctic XXI century. 2026;(3):5-21. (In Russ.) https://doi.org/10.25587/3034-7378-2026-14-3-5-21
JATS XML











