<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">arkhumsci</journal-id><journal-title-group><journal-title xml:lang="ru">Арктика XXI век</journal-title><trans-title-group xml:lang="en"><trans-title>Arctic XXI century</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">3034-7378</issn><issn pub-type="epub">3034-7386</issn><publisher><publisher-name>Северо-Восточный федеральный университет им. М. К. Аммосова</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.25587/3034-7378-2025-4-56-78</article-id><article-id custom-type="elpub" pub-id-type="custom">arkhumsci-242</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>Статьи</subject></subj-group></article-categories><title-group><article-title>Подходы к распознаванию и синтезу речи для якутского языка  на основе нейросетей архитектуры Transformer</article-title><trans-title-group xml:lang="en"><trans-title>Transformer-Based Neural Network Approaches for Speech Recognition and Synthesis in the Sakha Language</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-9445-2726</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Степанов</surname><given-names>С. П.</given-names></name><name name-style="western" xml:lang="en"><surname>Stepanov</surname><given-names>S. P.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Степанов Сергей Павлович – кандидат физико-математических наук, ведущий  научный  сотрудник,  руководитель, лаборатория «Вычислительные технологии и искусственныйинтеллект» , Институт математики и информатики</p><p>Якутск</p><p>WoS Researcher ID: F-7549-2017</p><p>Scopus Author ID: 56419440700</p><p>Elibrary Author ID: 856700</p></bio><bio xml:lang="en"><p>Sergei  P.  Stepanov  –  Cand.  Sci.  (Physics  and  Mathematics),  Head,  the Laboratory  “Computational  Technologies  and  Artificial  Intelligence”,  Institute of  Mathematics  and  Information  Science</p><p>Yakutsk</p><p>WoS  Researcher ID:  F-7549-2017</p><p>Scopus  Author  ID:  56419440700</p><p>Elibrary AuthorID: 856700</p></bio><email xlink:type="simple">sp.stepanov@s-vfu.ru</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-4688-7762</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Чжан</surname><given-names>Д.</given-names></name><name name-style="western" xml:lang="en"><surname>Zhang</surname><given-names>Dong</given-names></name></name-alternatives><bio xml:lang="ru"><p>Чжан Дун – кандидат физико-математических наук, доцент  </p><p> WoS  Researcher ID:  ACW-5232-2022</p><p>Scopus  Author ID: 57212194896</p><p>Цюйфу, Шаньдун</p></bio><bio xml:lang="en"><p>Dong  Zhang – Cand. Sci. (Physics and Mathematics), Associate Professor</p><p>Qufu,  Shandong</p><p>WoS  Researcher ID:  ACW-5232-2022</p><p>Scopus  Author ID: 57212194896</p></bio><email xlink:type="simple">dz_zhangdong@163.com</email><xref ref-type="aff" rid="aff-2"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Алексеева</surname><given-names>А. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Alekseeva</surname><given-names>A. A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Алексеева Алтана Александровна – лаборант, лаборатория «Вычислительные технологии и искусственныйинтеллект», Институт математики и информатики</p><p>Якутск</p></bio><bio xml:lang="en"><p>Altana A. Alekseeva –  Research  Assistant,     the Laboratory  “Computational  Technologies  and  Artificial  Intelligence”,  Institute of  Mathematics  and  Information  Science</p><p> Yakutsk</p></bio><email xlink:type="simple">altana.alexeeva@gmail.com</email><xref ref-type="aff" rid="aff-3"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Апросимов</surname><given-names>В. Л.</given-names></name><name name-style="western" xml:lang="en"><surname>Aprosimov</surname><given-names>V. L.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Апросимов Владислав Леонидович – лаборант, лаборатория «Вычислительные технологии и искусственный интеллект», Институт математики и информатики</p><p>Якутск</p></bio><bio xml:lang="en"><p>Vladislav L. Aprosimov – Research  Assistant,     the Laboratory  “Computational  Technologies  and  Artificial  Intelligence”,  Institute of  Mathematics  and  Information  Science</p><p> Yakutsk</p></bio><email xlink:type="simple">malysay88@gmail.com</email><xref ref-type="aff" rid="aff-3"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Федоров</surname><given-names>Дь. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Fedorov</surname><given-names>Dj. A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Федоров  Дьулуур  Андрианович  –  лаборант,  лаборатория  «Вычислительные технологии и искусственный интеллект», Институт математики и информатики</p><p>Якутск</p></bio><bio xml:lang="en"><p>Djuluur A.  Fedorov – Research  Assistant,     the Laboratory  “Computational  Technologies  and  Artificial  Intelligence”,  Institute of  Mathematics  and  Information  Science</p><p> Yakutsk</p></bio><email xlink:type="simple">fjuluur@mail.ru</email><xref ref-type="aff" rid="aff-3"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Леверьев</surname><given-names>В. С.</given-names></name><name name-style="western" xml:lang="en"><surname>Leveryev</surname><given-names>V. S.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Леверьев  Владимир  Семенович  –  лаборант,  лаборатория  «Вычислительные технологии и искусственный интеллект», Институт математики и информатики</p><p>Якутск </p></bio><bio xml:lang="en"><p>Vladimir S.  Leveryev – Research  Assistant,     the Laboratory  “Computational  Technologies  and  Artificial  Intelligence”,  Institute of  Mathematics  and  Information  Science</p><p> Yakutsk</p></bio><email xlink:type="simple">leverev.vs@svfu.ru</email><xref ref-type="aff" rid="aff-3"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Новгородов</surname><given-names>Т. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Novgorodov</surname><given-names>T. A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Новгородов  Туйгун  Александрович  –  лаборант,  лаборатория «Вычислительные технологии  и  искусственный  интеллект»,  Институт  математики  и  информатики</p><p>Якутск </p></bio><bio xml:lang="en"><p>Tuygun A. Novgorodov – Research  Assistant,     the Laboratory  “Computational  Technologies  and  Artificial  Intelligence”,  Institute of  Mathematics  and  Information  Science</p><p> Yakutsk</p></bio><email xlink:type="simple">tuygun2000@gmail.com</email><xref ref-type="aff" rid="aff-3"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Подорожная</surname><given-names>Е. С.</given-names></name><name name-style="western" xml:lang="en"><surname>Podorozhnaya</surname><given-names>E. S.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Подорожная Екатерина  Сергеевна  –  лаборант,  лаборатория  «Вычислительные технологии  и  искусственный  интеллект»,  Институт  математики  и  информатики</p><p>Якутск </p></bio><bio xml:lang="en"><p>Ekaterina S. Podorozhnaya – Research  Assistant,     the Laboratory  “Computational  Technologies  and  Artificial  Intelligence”,  Institute of  Mathematics  and  Information  Science</p><p> Yakutsk</p></bio><email xlink:type="simple">ekpodor@gmail.com</email><xref ref-type="aff" rid="aff-4"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Захаров</surname><given-names>Т. З.</given-names></name><name name-style="western" xml:lang="en"><surname>Zakharov</surname><given-names>T. Z.</given-names></name></name-alternatives><bio xml:lang="ru"><p>ЗахаровТимур Захарович – лаборант, лаборатория «Вычислительныетехнологии и искусственный интеллект», Институт математики и информатики</p><p> Якутск</p></bio><bio xml:lang="en"><p>Timur  Z.  Zakharov  –  Research  Assistant,    laboratory  “Computational Technologies and Artificial Intelligence”, Institute of Mathematics and Information Science</p><p>Yakutsk</p></bio><email xlink:type="simple">timu.zaxarov40@gmail.com</email><xref ref-type="aff" rid="aff-5"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Северо-Восточный федеральный университет им. М. К. Аммосова</institution><country>Россия</country></aff><aff xml:lang="en"><institution>M.K. Ammosov  North-Eastern  Federal University</institution><country>Russian Federation</country></aff></aff-alternatives><aff-alternatives id="aff-2"><aff xml:lang="ru"><institution>Цюйфуский педагогический университет</institution><country>Китай</country></aff><aff xml:lang="en"><institution>Qufu  Normal  University</institution><country>China</country></aff></aff-alternatives><aff-alternatives id="aff-3"><aff xml:lang="ru"><institution>Северо-Восточный федеральный университет им. М. К. Аммосова</institution><country>Россия</country></aff><aff xml:lang="en"><institution>M.K.  Ammosov  North-Eastern  Federal  University</institution><country>Russian Federation</country></aff></aff-alternatives><aff-alternatives id="aff-4"><aff xml:lang="ru"><institution>Северо-Восточный  федеральный  университет  им.  М.К.Аммосова</institution><country>Russian Federation</country></aff><aff xml:lang="en"><institution>M.K.  Ammosov  North-Eastern  Federal  University</institution><country>Russian Federation</country></aff></aff-alternatives><aff-alternatives id="aff-5"><aff xml:lang="ru"><institution>Северо-Восточный федеральный университет им. М.К. Аммосова</institution><country>Russian Federation</country></aff><aff xml:lang="en"><institution>M.K.  Ammosov  North-Eastern  Federal  University</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2025</year></pub-date><pub-date pub-type="epub"><day>17</day><month>01</month><year>2026</year></pub-date><volume>0</volume><issue>4</issue><fpage>56</fpage><lpage>78</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Степанов С.П., Чжан Д., Алексеева А.А., Апросимов В.Л., Федоров Д.А., Леверьев В.С., Новгородов Т.А., Подорожная Е.С., Захаров Т.З., 2025</copyright-statement><copyright-year>2025</copyright-year><copyright-holder xml:lang="ru">Степанов С.П., Чжан Д., Алексеева А.А., Апросимов В.Л., Федоров Д.А., Леверьев В.С., Новгородов Т.А., Подорожная Е.С., Захаров Т.З.</copyright-holder><copyright-holder xml:lang="en">Stepanov S.P., Zhang D., Alekseeva A.A., Aprosimov V.L., Fedorov D.A., Leveryev V.S., Novgorodov T.A., Podorozhnaya E.S., Zakharov T.Z.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://www.arcticjournal.ru/jour/article/view/242">https://www.arcticjournal.ru/jour/article/view/242</self-uri><abstract><p>Новейшие достижения в области искусственного интеллекта и глубокого обучения  кардинально  преобразовали  ландшафт  технологий  обработки  устной речи.  Автоматическое  распознавание  речи  (ASR)  и  синтез  речи  (TTS)  стали ключевыми  компонентами,  обеспечивающими  цифровую  доступность  для различных  языковых  сообществ.  Якутский  язык,  представляющий  северо-восточную ветвь тюркской языковой семьи, продолжает сталкиваться со значительными  технологическими  барьерами,  вызванными  недостаточностью цифровых  ресурсов,  ограниченностью  размеченных  корпусов  и  отсутствием готовых к промышленному использованию систем обработки речи. В данном комплексном  исследовании  изучается  целесообразность  и  эффективность адаптации  современных  нейросетевых  архитектур  на  основе  трансформеров для задач двунаправленного речевого преобразования в якутском языке. Наша работа включает детальный анализ encoder-decoder моделей, а именно: Whisper-large-v3  от  OpenAI  и  Wav2Vec2-BERT  от  Meta  для  преобразования  голоса  в текст, а также системы XTTS-v2 от Coqui для генерации речи из текста. Особое внимание уделяется решению лингвистических и технических проблем, присущих якутскому языку, включая его сложную агглютинативную морфологическую структуру, системные законы сингармонизма и уникальный фонемный состав,  содержащий  звуки,  отсутствующие  в  большинстве  индоевропейских языков. Экспериментальная оценка показывает, что полное дообучение модели Whisper-large-v3 обеспечивает исключительно высокую точность распознавания с коэффициентом ошибок по словам (WER) 8%, в то время как самообучаемая архитектура Wav2Vec2-BERT достигает WER 13% при использовании статистического n-граммного языкового моделирования. Нейросетевая система  синтеза  демонстрирует  устойчивую  производительность  даже  при  ограниченном объеме обучающих данных, достигая среднего значения функции потерь 2,49  после  длительной  оптимизации  обучения  и  практического  развертывания через бот в мессенджере Telegram. Кроме того, ансамблевый мета-стэкинг, объединяющий обе архитектуры распознавания, позволяет достичь WER 27%, что доказывает их эффективную взаимодополняемость через арбитраж гипотез. Полученные результаты подтверждают, что методы трансферного обучения представляют собой жизнеспособный путь для создания речевых технологий, обслуживающих цифрово недостаточно представленные языковые сообщества.</p></abstract><trans-abstract xml:lang="en"><p>Recent breakthroughs in artificial intelligence and deep learning have fundamentally transformed the landscape of spoken language processing technologies. Automatic speech recognition (ASR) and text-to-speech (TTS) synthesis have emerged as essential components driving digital accessibility across diverse linguistic communities. The Sakha language, representing the northeastern branch of the Turkic language family, continues to face substantial technological barriers stemming from insufficient digital resources,  limited  annotated  corpora,  and  the  absence  of  production-ready  speech processing systems. This comprehensive investigation examines the feasibility and effectiveness of adapting contemporary transformer-based neural architectures for bidirectional speech conversion tasks in Sakha. Our research encompasses detailed analysis  of  encoder-decoder  frameworks,  specifically  OpenAI’s  Whisper  large-v3 and  Meta’s  Wav2Vec2-BERT  for  voice-to-text  transformation,  alongside  Coqui’s XTTS-v2  system  for  text-to-voice  generation.  Particular  emphasis  is  placed  on addressing linguistic and technical obstacles inherent to Sakha, including its complex agglutinative  morphological  structure,  systematic  vowel  harmony  patterns,  and distinctive phonemic inventory featuring sounds absent from most Indo-European languages.  Experimental  evaluation  demonstrates  that  comprehensive  fine-tuning of  Whisper-large-v3  achieves  exceptional  recognition  accuracy  with  word  error rate (WER) of 8%, while the self-supervised Wav2Vec2-BERT architecture attains 13%  WER  when  augmented  with  statistical  n-gram  language  modeling.  The neural synthesis system exhibits robust performance despite minimal training data availability, achieving average loss of 2.49 following extended training optimization and practical deployment via Telegram messaging bot. Additionally, ensemble meta-stacking combining both recognition architectures achieves 27% WER, demonstrating effective  complementarity  through  learned  hypothesis  arbitration.  These  findings validate transfer learning methodologies as viable pathways for developing speech technologies serving digitally underrepresented linguistic communities.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>якутский язык</kwd><kwd>автоматическое распознавание речи</kwd><kwd>синтез речи из текста</kwd><kwd>нейронные сети</kwd><kwd>Whisper</kwd><kwd>Wav2Vec2-BERT</kwd><kwd>Coqui XTTS-v2</kwd><kwd>архитектура Transformer</kwd><kwd>малоресурсные языки</kwd><kwd>трансферное обучение</kwd><kwd>агглютинативная морфология</kwd></kwd-group><kwd-group xml:lang="en"><kwd>Sakha language</kwd><kwd>automatic speech recognition</kwd><kwd>text-to-speech synthesis</kwd><kwd>neural  networks</kwd><kwd>Whisper</kwd><kwd>Wav2Vec2-BERT</kwd><kwd>Coqui  XTTS-v2</kwd><kwd>transformer architecture</kwd><kwd>low-resource languages</kwd><kwd>transfer learning</kwd><kwd>agglutinative morphology</kwd></kwd-group><funding-group><funding-statement xml:lang="ru">Данное исследование было выполнено при финансовой поддержке  Лаборатории  искусственного  интеллекта  Республики  Саха  (Якутия), а также при финансовой поддержке РНФ в рамках реализации проекта «Языки и культуры народов Севера и Арктики РФ: комплексные социогуманитарные исследования (на основе анализа больших данных)» № 25-78-30006 от 22.05.2025 г.</funding-statement><funding-statement xml:lang="en">This  research  was  conducted  with  financial  support  from  the Artificial  Intelligence Laboratory of the Republic of Sakha (Yakutia) and with the financial  support of the Russian Science Foundation “Languages and Cultures of the Peoples  of  the  North  and  the  Arctic  of  the  Russian  Federation:  Comprehensive  socio humanitarian research (on the basis of big data)” No 25-78-30006 (22.05.2025)</funding-statement></funding-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Besacier L, Barnard E, Karpov A, Schultz T. Automatic speech recognition for under-resourced languages: A survey. Speech Communication. 2014;(56): 85–100. DOI: https://doi.org/10.1016/j.specom.2013.07.008</mixed-citation><mixed-citation xml:lang="en">Besacier L, Barnard E, Karpov A, Schultz T. Automatic speech recognition for under-resourced languages: A survey. Speech Communication. 2014;(56): 85–100. DOI: https://doi.org/10.1016/j.specom.2013.07.008</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Joshi P, Santy S, Buber A, Bali K, Choudhury M. The state and fate of linguistic diversity and inclusion in the NLP world. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020: 6282–6293.</mixed-citation><mixed-citation xml:lang="en">Joshi P, Santy S, Buber A, Bali K, Choudhury M. The state and fate of linguistic diversity and inclusion in the NLP world. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020: 6282–6293.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Pakendorf B. Contact in the prehistory of the Sakha (Yakuts): Linguistic and genetic perspectives. LOT Publications: Utrecht. 2007</mixed-citation><mixed-citation xml:lang="en">Pakendorf B. Contact in the prehistory of the Sakha (Yakuts): Linguistic and genetic perspectives. LOT Publications: Utrecht. 2007</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Johanson L, Csató ÉA. The Turkic Languages. Routledge Language Family Series. Routledge: London. 2021. DOI: https://doi.org/10.4324/9781003243809</mixed-citation><mixed-citation xml:lang="en">Johanson L, Csató ÉA. The Turkic Languages. Routledge Language Family Series. Routledge: London. 2021. DOI: https://doi.org/10.4324/9781003243809</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Dyachkovsky ND. Sound structure of the Yakut language. Part 1: Vocalism (Дьячковский Н.Д. Звуковой строй якутского языка. Вокализм). Yakutsk: Yakutsk publishing house. 1971 (in Russian).</mixed-citation><mixed-citation xml:lang="en">Dyachkovsky ND. Sound structure of the Yakut language. Part 1: Vocalism (Дьячковский Н.Д. Звуковой строй якутского языка. Вокализм). Yakutsk: Yakutsk publishing house. 1971 (in Russian).</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Dyachkovsky ND. Sound structure of the Yakut language. Part 2: Consonantism (Дьячковский Н.Д. Звуковой строй якутского языка. Консонантизм). Yakutsk: Yakutsk publishing house. 1977 (in Russian).</mixed-citation><mixed-citation xml:lang="en">Dyachkovsky ND. Sound structure of the Yakut language. Part 2: Consonantism (Дьячковский Н.Д. Звуковой строй якутского языка. Консонантизм). Yakutsk: Yakutsk publishing house. 1977 (in Russian).</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Mussakhojayeva S, Dauletbek K, Yeshpanov R, Varol HA. Multilingual speech recognition for Turkic languages. Information. 2023;14(2): 74. DOI: https://doi.org/10.3390/info14020074</mixed-citation><mixed-citation xml:lang="en">Mussakhojayeva S, Dauletbek K, Yeshpanov R, Varol HA. Multilingual speech recognition for Turkic languages. Information. 2023;14(2): 74. DOI: https://doi.org/10.3390/info14020074</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Rabiner LR. A tutorial on hidden Markov models and selected applications in speech recognition. Proceedings of the IEEE. 1989;77(2):257–286. DOI: http://dx.doi.org/10.1109/5.18626</mixed-citation><mixed-citation xml:lang="en">Rabiner LR. A tutorial on hidden Markov models and selected applications in speech recognition. Proceedings of the IEEE. 1989;77(2):257–286. DOI: http://dx.doi.org/10.1109/5.18626</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Graves A, Mohamed A, Hinton G. Speech recognition with deep recurrent neural networks. Proceedings of International Conference on Acoustics, Speech and Signal Processing. 2013:6645–6649. DOI: https://doi.org/10.1109/ICASSP.2013.6638947</mixed-citation><mixed-citation xml:lang="en">Graves A, Mohamed A, Hinton G. Speech recognition with deep recurrent neural networks. Proceedings of International Conference on Acoustics, Speech and Signal Processing. 2013:6645–6649. DOI: https://doi.org/10.1109/ICASSP.2013.6638947</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Conneau A, Baevski A, Collobert R, Mohamed A, Auli M. Unsupervised cross-lingual representation learning for speech recognition. In Proceedings of Interspeech. 2021:2426–2430. DOI: https://doi.org/10.48550/arXiv.2006.13979</mixed-citation><mixed-citation xml:lang="en">Conneau A, Baevski A, Collobert R, Mohamed A, Auli M. Unsupervised cross-lingual representation learning for speech recognition. In Proceedings of Interspeech. 2021:2426–2430. DOI: https://doi.org/10.48550/arXiv.2006.13979</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017:5998–6008. DOI: https://doi.org/10.48550/arXiv.1706.03762</mixed-citation><mixed-citation xml:lang="en">Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017:5998–6008. DOI: https://doi.org/10.48550/arXiv.1706.03762</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Radford A, Kim JW, Xu T, Brockman G, McLeavey C, Sutskever I. Robust speech recognition via large-scale weak supervision. In Proceedings of International Conference on Machine Learning. 2023:28492–28518. DOI: https://doi.org/10.48550/arXiv.2212.04356</mixed-citation><mixed-citation xml:lang="en">Radford A, Kim JW, Xu T, Brockman G, McLeavey C, Sutskever I. Robust speech recognition via large-scale weak supervision. In Proceedings of International Conference on Machine Learning. 2023:28492–28518. DOI: https://doi.org/10.48550/arXiv.2212.04356</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Du W, Maimaitiyiming Y, Nijat M, Li L, Hamdulla A, Wang D. Automatic speech recognition for Uyghur, Kazakh, and Kyrgyz: An overview. Applied Sciences. 2023;13(1):326. DOI: https://doi.org/10.3390/app13010326</mixed-citation><mixed-citation xml:lang="en">Du W, Maimaitiyiming Y, Nijat M, Li L, Hamdulla A, Wang D. Automatic speech recognition for Uyghur, Kazakh, and Kyrgyz: An overview. Applied Sciences. 2023;13(1):326. DOI: https://doi.org/10.3390/app13010326</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Yeshpanov R, Mussakhojayeva S, Khassanov Y. Multilingual text-tospeech synthesis for Turkic languages using transliteration. In Proceedings of Interspeech. 2023:5521–5525. DOI: https://doi.org/10.48550/arXiv.2305.15749</mixed-citation><mixed-citation xml:lang="en">Yeshpanov R, Mussakhojayeva S, Khassanov Y. Multilingual text-tospeech synthesis for Turkic languages using transliteration. In Proceedings of Interspeech. 2023:5521–5525. DOI: https://doi.org/10.48550/arXiv.2305.15749</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Kim J, Kim S, Kong J, Yoon S. Glow-TTS: A generative low for text-to-speech via monotonic alignment search. In Proceedings of the International Conference on Neural Information Processing Systems. 2020:8067–8077. DOI: https://doi.org/10.48550/arXiv.2005.11129</mixed-citation><mixed-citation xml:lang="en">Kim J, Kim S, Kong J, Yoon S. Glow-TTS: A generative low for text-to-speech via monotonic alignment search. In Proceedings of the International Conference on Neural Information Processing Systems. 2020:8067–8077. DOI: https://doi.org/10.48550/arXiv.2005.11129</mixed-citation></citation-alternatives></ref><ref id="cit16"><label>16</label><citation-alternatives><mixed-citation xml:lang="ru">Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In Proceedings of the International Conference on Machine Learning. 2021:5530–5540. DOI: https://doi.org/10.48550/arXiv.2106.06103</mixed-citation><mixed-citation xml:lang="en">Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In Proceedings of the International Conference on Machine Learning. 2021:5530–5540. DOI: https://doi.org/10.48550/arXiv.2106.06103</mixed-citation></citation-alternatives></ref><ref id="cit17"><label>17</label><citation-alternatives><mixed-citation xml:lang="ru">Kong J, Kim J, Bae J. HiFi-GAN: Generative adversarial networks for eficient and high idelity speech synthesis. In Proceedings of the International Conference on Neural Information Processing Systems. 2020:17022–17033. DOI: https://doi.org/10.48550/arXiv.2010.05646</mixed-citation><mixed-citation xml:lang="en">Kong J, Kim J, Bae J. HiFi-GAN: Generative adversarial networks for eficient and high idelity speech synthesis. In Proceedings of the International Conference on Neural Information Processing Systems. 2020:17022–17033. DOI: https://doi.org/10.48550/arXiv.2010.05646</mixed-citation></citation-alternatives></ref><ref id="cit18"><label>18</label><citation-alternatives><mixed-citation xml:lang="ru">Shen J, Pang R, Weiss RJ, Schuster M, Jaitly N, Yang Z, Chen Z, Zhang Y, Wang Y, Skerrv-Ryan R., et al. Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions. In Proceedings of IEEE ICASSP. 2018:4779–4783. DOI: https://doi.org/10.48550/arXiv.1712.05884</mixed-citation><mixed-citation xml:lang="en">Shen J, Pang R, Weiss RJ, Schuster M, Jaitly N, Yang Z, Chen Z, Zhang Y, Wang Y, Skerrv-Ryan R., et al. Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions. In Proceedings of IEEE ICASSP. 2018:4779–4783. DOI: https://doi.org/10.48550/arXiv.1712.05884</mixed-citation></citation-alternatives></ref><ref id="cit19"><label>19</label><citation-alternatives><mixed-citation xml:lang="ru">Ren Y, Hu C, Tan X, Qin T, Zhao S, Zhao Z, Liu T-Y. FastSpeech 2: Fast and high-quality end-to-end text to speech. In Proceedings of ICLR. 2021. DOI: https://doi.org/10.48550/arXiv.2006.04558</mixed-citation><mixed-citation xml:lang="en">Ren Y, Hu C, Tan X, Qin T, Zhao S, Zhao Z, Liu T-Y. FastSpeech 2: Fast and high-quality end-to-end text to speech. In Proceedings of ICLR. 2021. DOI: https://doi.org/10.48550/arXiv.2006.04558</mixed-citation></citation-alternatives></ref><ref id="cit20"><label>20</label><citation-alternatives><mixed-citation xml:lang="ru">Karibayeva A, Karyukin V, Abduali B, Amirova D. Speech recognition and synthesis models and platforms for the Kazakh language. Information. 2025;16(10):879. DOI: https://doi.org/10.3390/info16100879</mixed-citation><mixed-citation xml:lang="en">Karibayeva A, Karyukin V, Abduali B, Amirova D. Speech recognition and synthesis models and platforms for the Kazakh language. Information. 2025;16(10):879. DOI: https://doi.org/10.3390/info16100879</mixed-citation></citation-alternatives></ref><ref id="cit21"><label>21</label><citation-alternatives><mixed-citation xml:lang="ru">Ardila R, Branson M, Davis K, Kohler M, Meyer J, Henretty M, Morais R, Saunders L, Tyers F, Weber G. Common Voice: A massively-multilingual speech corpus. In Proceedings of LREC. 2020:4218–4222. DOI: https://doi.org/10.48550/arXiv.1912.06670</mixed-citation><mixed-citation xml:lang="en">Ardila R, Branson M, Davis K, Kohler M, Meyer J, Henretty M, Morais R, Saunders L, Tyers F, Weber G. Common Voice: A massively-multilingual speech corpus. In Proceedings of LREC. 2020:4218–4222. DOI: https://doi.org/10.48550/arXiv.1912.06670</mixed-citation></citation-alternatives></ref><ref id="cit22"><label>22</label><citation-alternatives><mixed-citation xml:lang="ru">Park DS, Chan W, Zhang Y, Chiu C-C, Zoph B, Cubuk ED, Le QV. SpecAugment: A simple data augmentation method for automatic speech recognition. In Proceedings of Interspeech. 2019:2613–2617. DOI: https://doi.org/10.21437/Interspeech.2019-2680</mixed-citation><mixed-citation xml:lang="en">Park DS, Chan W, Zhang Y, Chiu C-C, Zoph B, Cubuk ED, Le QV. SpecAugment: A simple data augmentation method for automatic speech recognition. In Proceedings of Interspeech. 2019:2613–2617. DOI: https://doi.org/10.21437/Interspeech.2019-2680</mixed-citation></citation-alternatives></ref><ref id="cit23"><label>23</label><citation-alternatives><mixed-citation xml:lang="ru">Kudo T, Richardson J. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of EMNLP System Demonstrations. 2018:66–71. DOI: https://doi.org/10.48550/arXiv.1808.06226</mixed-citation><mixed-citation xml:lang="en">Kudo T, Richardson J. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of EMNLP System Demonstrations. 2018:66–71. DOI: https://doi.org/10.48550/arXiv.1808.06226</mixed-citation></citation-alternatives></ref><ref id="cit24"><label>24</label><citation-alternatives><mixed-citation xml:lang="ru">Baevski A, Zhou Y, Mohamed A, Auli M. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the 34th International Conference on Neural Information Processing System. 2020:12449-12460. DOI: https://doi.org/10.48550/arXiv.2006.11477</mixed-citation><mixed-citation xml:lang="en">Baevski A, Zhou Y, Mohamed A, Auli M. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the 34th International Conference on Neural Information Processing System. 2020:12449-12460. DOI: https://doi.org/10.48550/arXiv.2006.11477</mixed-citation></citation-alternatives></ref><ref id="cit25"><label>25</label><citation-alternatives><mixed-citation xml:lang="ru">Graves A, Fernández S, Gomez F, Schmidhuber J. Connectionist temporal classiication: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of ICML. 2006:369–376. DOI: https://doi.org/10.1145/1143844.1143891</mixed-citation><mixed-citation xml:lang="en">Graves A, Fernández S, Gomez F, Schmidhuber J. Connectionist temporal classiication: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of ICML. 2006:369–376. DOI: https://doi.org/10.1145/1143844.1143891</mixed-citation></citation-alternatives></ref><ref id="cit26"><label>26</label><citation-alternatives><mixed-citation xml:lang="ru">Casanova E, Weber J, Shulby C, Junior AC, Gölge E, Müller MA. XTTS: A massively multilingual zero-shot text-to-speech model. In Proceedings of Interspeech. 2024. DOI: https://doi.org/10.48550/arXiv.2406.04904</mixed-citation><mixed-citation xml:lang="en">Casanova E, Weber J, Shulby C, Junior AC, Gölge E, Müller MA. XTTS: A massively multilingual zero-shot text-to-speech model. In Proceedings of Interspeech. 2024. DOI: https://doi.org/10.48550/arXiv.2406.04904</mixed-citation></citation-alternatives></ref><ref id="cit27"><label>27</label><citation-alternatives><mixed-citation xml:lang="ru">Panayotov V, Chen G, Povey D, Khudanpur S. LibriSpeech: An ASR corpus based on public domain audio books. In Proceedings of IEEE ICASSP. 2015:5206-5210. DOI: https://doi.org/10.1109/ICASSP.2015.7178964</mixed-citation><mixed-citation xml:lang="en">Panayotov V, Chen G, Povey D, Khudanpur S. LibriSpeech: An ASR corpus based on public domain audio books. In Proceedings of IEEE ICASSP. 2015:5206-5210. DOI: https://doi.org/10.1109/ICASSP.2015.7178964</mixed-citation></citation-alternatives></ref><ref id="cit28"><label>28</label><citation-alternatives><mixed-citation xml:lang="ru">Dyakonov AG. Ensembles in machine learning: Methods and applications. Data Science Course Materials (Дьяконов А.Г. Ансамбли в машинном обучении). 2019. Available at: https://alexanderdyakonov.wordpress.com (accessed 07.09.2025).</mixed-citation><mixed-citation xml:lang="en">Dyakonov AG. Ensembles in machine learning: Methods and applications. Data Science Course Materials (Дьяконов А.Г. Ансамбли в машинном обучении). 2019. Available at: https://alexanderdyakonov.wordpress.com (accessed 07.09.2025).</mixed-citation></citation-alternatives></ref><ref id="cit29"><label>29</label><citation-alternatives><mixed-citation xml:lang="ru">Wolpert DH. Stacked generalization. Neural Networks 1992;5(2): 241-259. DOI: https://doi.org/10.1016/S0893-6080(05)80023-1</mixed-citation><mixed-citation xml:lang="en">Wolpert DH. Stacked generalization. Neural Networks 1992;5(2): 241-259. DOI: https://doi.org/10.1016/S0893-6080(05)80023-1</mixed-citation></citation-alternatives></ref><ref id="cit30"><label>30</label><citation-alternatives><mixed-citation xml:lang="ru">Breiman L. Stacked regressions. Machine Learning. 1996;24(1):49–64.</mixed-citation><mixed-citation xml:lang="en">Breiman L. Stacked regressions. Machine Learning. 1996;24(1):49–64.</mixed-citation></citation-alternatives></ref><ref id="cit31"><label>31</label><citation-alternatives><mixed-citation xml:lang="ru">Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R, Funtowicz M, et al. Transformers: State-of-the-art natural language processing. In Proceedings of the Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 2020:38–45. DOI: https://doi.org/10.18653/v1/2020.emnlp-demos.6</mixed-citation><mixed-citation xml:lang="en">Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R, Funtowicz M, et al. Transformers: State-of-the-art natural language processing. In Proceedings of the Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 2020:38–45. DOI: https://doi.org/10.18653/v1/2020.emnlp-demos.6</mixed-citation></citation-alternatives></ref><ref id="cit32"><label>32</label><citation-alternatives><mixed-citation xml:lang="ru">Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L, et al. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the International Conference on Neural Information Processing Systems. 2019:8026–8037. DOI: https://doi.org/10.48550/arXiv.1912.01703</mixed-citation><mixed-citation xml:lang="en">Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L, et al. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the International Conference on Neural Information Processing Systems. 2019:8026–8037. DOI: https://doi.org/10.48550/arXiv.1912.01703</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
