<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">altaistika</journal-id><journal-title-group><journal-title xml:lang="ru">Алтаистика. Altaistics</journal-title><trans-title-group xml:lang="en"><trans-title>Altaistics</trans-title></trans-title-group></journal-title-group><issn pub-type="epub">2782-6627</issn><publisher><publisher-name>Северо-Восточный федеральный университет имени М.К. Аммосова</publisher-name></publisher></journal-meta><article-meta><article-id custom-type="elpub" pub-id-type="custom">altaistika-29</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>Статьи</subject></subj-group></article-categories><title-group><article-title>О разработке лингвистической базы данных онтологического типа, как ресурса для лингвопроцессоров</article-title><trans-title-group xml:lang="en"><trans-title>On development of ontological linguistic database as a resource for language processors</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Гатиатуллин</surname><given-names>А. Р.</given-names></name><name name-style="western" xml:lang="en"><surname>Gatiatullin</surname><given-names>A. R.</given-names></name></name-alternatives><bio xml:lang="ru"><p>ГАТИАТУЛЛИН Айрат Рафизович – к. тех. н., ведущий научный сотрудник</p><p>г. Казань</p></bio><bio xml:lang="en"><p>GATIATULLIN Ayrat Rafizovich – Candidate of Technical Sciences, Leading Researcher</p><p>Kazan</p></bio><email xlink:type="simple">ayrat.gatiatullin@gmail.com</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Прокопьев</surname><given-names>Н. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Prokopiev</surname><given-names>N. A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>ПРОКОПЬЕВ Николай Аркадьевич – научный сотрудник</p><p>г. Казань</p></bio><bio xml:lang="en"><p>PROKOPYEV Nikolay Arkadievich – Researcher</p><p>Kazan</p></bio><email xlink:type="simple">nikolai.prokopyev@gmail.com</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Академия наук Республики Татарстан</institution><country>Россия</country></aff><aff xml:lang="en"><institution>Tatarstan Academy of Sciences</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2021</year></pub-date><pub-date pub-type="epub"><day>27</day><month>02</month><year>2022</year></pub-date><volume>1</volume><issue>1</issue><fpage>77</fpage><lpage>88</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Гатиатуллин А.Р., Прокопьев Н.А., 2022</copyright-statement><copyright-year>2022</copyright-year><copyright-holder xml:lang="ru">Гатиатуллин А.Р., Прокопьев Н.А.</copyright-holder><copyright-holder xml:lang="en">Gatiatullin A.R., Prokopiev N.A.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://www.altaisticsvfu.ru/jour/article/view/29">https://www.altaisticsvfu.ru/jour/article/view/29</self-uri><abstract><p>В данной статье описывается разработка лингвистических онтологических баз данных для тюркских языков, которые могут быть использованы в целом ряде лингвистических процессоров для обработки текстов на тюркских языках. Актуальность данной работы заключается в том, что несмотря на активные разработки для тюркских языков в последние 10-15 лет практически все тюркские языки (кроме турецкого) продолжают относиться к типу малоресурсных языков. Это связано с тем, что для тюркских языков наблюдается дефицит лингвистических ресурсов, применимых в различных компьютерных разработках по обработке естественного языка. Это могут быть разного рода онтологические базы данных типа WordNet, FrameNet, VerbNet, РуТез и др., а также комбинации этих ресурсов с электронными корпусами. Подобные онтологические базы данных могут быть использованы в различных информационно-справочных системах, при создании синтаксических, семантических и семантико-синтаксических анализаторов, а также учебных и научных прикладных программ. В предлагаемой нами работе представлен подход, который объединяет онтологические модели фреймового и таксономического типа, структурно-параметрическую модель тюркской морфемы в единую интегральную модель. В основу разработки такой модели изначально положены принципы многоязычности, многофункциональности и прагматической ориентированности. Многоязычность предполагает универсальность для всех языков тюркской группы, а прагматическая ориентированность именно ориентированность на структурно-функциональные особенности языков агглютинативного типа. Создание программного обеспечения кроме вышеперечисленных теоретико-лингвистических методов и технологий предполагает использование технологий проектирования сложных баз данных, веб-программирования, клиентсерверных технологий. На основе интегральной онтологической модели создается многоязычная база данных для тюркских языков, которая используется для генерации правил контекстно-свободной грамматики и создания семантико-синтаксического анализатора. На вход данного анализатора поступают предложения на тюркских языках, а на выходе получаются структурированные данные. Получаемый таким образом анализатор применим для семантико-синтаксической разметки тюркских электронных корпусов и создания программ семантического поиска.</p></abstract><trans-abstract xml:lang="en"><p>This article describes the development of linguistic ontological database for Turkic languages, which can be used in a number of linguistic processors for texts processing in Turkology. The relevance of this work lies in the fact that despite active research and development for Turkic languages in the past 10-15 years, almost all (except for Turkish) continue to belong to the type of low-resource languages. This is due to the fact that there is a shortage of linguistic resources for Turkic languages applicable in software development for natural language processing. These can be various types of ontological databases such as WordNet, FrameNet, VerbNet, RuTez, etc., as well as combinations of these resources with electronic corpuses. Such ontological databases can be used in information and reference systems, in creation of syntactic, semantic and semanticsyntactic analyzers, as well as in educational and scientific applications. In our work, an integral approach is presented that combines ontological models of frame and taxonomic types, in a structural-parametric model of the Turkic Morpheme. The integral model is initially based on principles of multilingualism, multifunctionality and pragmatic orientation. Multilingualism presupposes universality for all languages of the Turkic group, and pragmatic orientation is in focus on structural and functional features of agglutinative languages. Creation of software, in addition to above methods and technologies, involves usage of technologies for designing complex databases, web programming, client-server technologies. The resulting database can be used to generate contextfree grammar rules for a semantic-syntactic analyzer, which receives sentences in Turkic languages as input, and produces structured data at output. The analyzer obtained in this way is applicable for semantic-syntactic tagging of Turkic electronic corpuses and for development of semantic search software.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>онтологические модели</kwd><kwd>семантический фрейм</kwd><kwd>тюркские языки</kwd></kwd-group><kwd-group xml:lang="en"><kwd>ontological models</kwd><kwd>semantic frame</kwd><kwd>Turkic languages</kwd></kwd-group><funding-group><funding-statement xml:lang="ru">Работа выполнена при поддержке гранта РФФИ 18-47-160014 «Разработка интегральной компьютерной модели и программного инструментария для семантико-синтаксического анализа татарских текстов»</funding-statement></funding-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Дыбо А.В., Шеймович А.В. (2014) Автоматический морфологический анализ для корпусов тюркских языков. Филология и культура, №2, с. 20-26.</mixed-citation><mixed-citation xml:lang="en">Dybo A.V., Shejmovich A.V. (2014) Avtomaticheskij morfologicheskij analiz dlja korpusov tjurkskih jazykov. Filologija i kul’tura, №2, s. 20-26.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Желтов П.В. (2002) Морфологический анализатор чувашского языка. Материалы Международной конференции студентов и аспирантов по фундаментальным наукам «Ломоносов 2002».</mixed-citation><mixed-citation xml:lang="en">Zheltov P.V. (2002) Morfologicheskij analizator chuvashskogo jazyka. Materialy Mezhdunarodnoj konferencii studentov i aspirantov po fundamental’nym naukam «Lomonosov 2002».</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Шарипбаев А.А., Бекманова Г.Т., Ергеш Б.Ж., Бурибаева А.К., Карабалаева М.Х. (2012) Интеллектуальный морфологический анализатор, основанный на семантических сетях. Материалы международной научно-технической конференции «Открытые семантические технологии проектирования интеллектуальных систем» (OSTIS-2012), с. 397-400.</mixed-citation><mixed-citation xml:lang="en">Sharipbaev A.A., Bekmanova G.T., Ergesh B.Zh., Buribaeva A.K., Karabalaeva M.H. (2012) Intellektual’nyj morfologicheskij analizator, osnovannyj na semanticheskih setjah. Materialy mezhdunarodnoj nauchno-tehnicheskoj konferencii «Otkrytye semanticheskie tehnologii proektirovanija intellektual’nyh sistem» (OSTIS-2012), s. 397-400.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Sharipbay A.A., Bekmanova G., Yergesh B., Mukanova A. (2014) Synchronized liner tree for morphological analysis and generation of the Kazakh language. Proceedings of the international conference “Turkic languages processing”, TurkLang 2014, pp. 113-117.</mixed-citation><mixed-citation xml:lang="en">Sharipbay A.A., Bekmanova G., Yergesh B., Mukanova A. (2014) Synchronized liner tree for morphological analysis and generation of the Kazakh language. Proceedings of the international conference “Turkic languages processing”, TurkLang 2014, pp. 113-117.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Orhun, M., Tantuğ A.C., Adalı E. (2010) Morphological Disambiguation Rules For Uyghur Language. IEEE International Conference on Software Engineering and Service Sciences (ICSESS), pp. 542-546. doi: 10.1109/ ICSESS.2010.5552304</mixed-citation><mixed-citation xml:lang="en">Orhun, M., Tantuğ A.C., Adalı E. (2010) Morphological Disambiguation Rules For Uyghur Language. IEEE International Conference on Software Engineering and Service Sciences (ICSESS), pp. 542-546. doi: 10.1109/ ICSESS.2010.5552304</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Sahin G.G., Adalı E. (2018) Annotation of semantic roles for the Turkish proposition bank, 52(3), pp. 673-706. doi: 10.1007/s10579-017-9390-y</mixed-citation><mixed-citation xml:lang="en">Sahin G.G., Adalı E. (2018) Annotation of semantic roles for the Turkish proposition bank, 52(3), pp. 673-706. doi: 10.1007/s10579-017-9390-y</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Eryiğit G., Nivre J., Oﬂazer K. (2008) Dependency Parsing of Turkish. Computational Linguistics, 34(3), pp. 357-389. doi: 10.1162/coli.2008.34.4.627</mixed-citation><mixed-citation xml:lang="en">Eryiğit G., Nivre J., Oflazer K. (2008) Dependency Parsing of Turkish. Computational Linguistics, 34(3), pp. 357-389. doi: 10.1162/coli.2008.34.4.627</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Lyashevskaya O., Kashkin E. (2015) FrameBank: A Database of Russian Lexical Constructions. Proceedings of the 4th International Conference on Analysis of Images, Social Networks and Texts (AIST 2015). Communications in Computer and Information Science, vol. 542, pp. 350-360. doi:10.1007/978-3-319-2</mixed-citation><mixed-citation xml:lang="en">Lyashevskaya O., Kashkin E. (2015) FrameBank: A Database of Russian Lexical Constructions. Proceedings of the 4th International Conference on Analysis of Images, Social Networks and Texts (AIST 2015). Communications in Computer and Information Science, vol. 542, pp. 350-360. doi:10.1007/978-3-319-2</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Turkish National Corpus (TNC). URL: http:// www.tnc.org.tr.</mixed-citation><mixed-citation xml:lang="en">Turkish National Corpus (TNC). URL: http:// www.tnc.org.tr.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Алматинский корпус казахского языка. URL: http://web-corpora.net/KazakhCorpus/search/.</mixed-citation><mixed-citation xml:lang="en">Almatinskij korpus kazahskogo jazyka. URL: http://web-corpora.net/KazakhCorpus/search/.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Корпус алтайского языка. URL: http://altay 2.gasu.ru.</mixed-citation><mixed-citation xml:lang="en">Korpus altajskogo jazyka. URL: http://altay 2.gasu.ru.</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Национальный корпус башкирского языка. URL: http://bashcorpus.ru.</mixed-citation><mixed-citation xml:lang="en">Nacional’nyj korpus bashkirskogo jazyka. URL: http://bashcorpus.ru.</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Башкирский поэтический корпус. URL: http:// web-corpora.net/bashcorpus/search/.</mixed-citation><mixed-citation xml:lang="en">Bashkirskij pojeticheskij korpus. URL: http:// web-corpora.net/bashcorpus/search/.</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Корпус татарского языка ‘Туган тел’. URL: http://tugantel.tatar.</mixed-citation><mixed-citation xml:lang="en">Korpus tatarskogo jazyka ‘Tugan tel’. URL: http://tugantel.tatar.</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Письменный корпус татарского языка. URL: http://www.corpus.tatar.</mixed-citation><mixed-citation xml:lang="en">Pis’mennyj korpus tatarskogo jazyka. URL: http://www.corpus.tatar.</mixed-citation></citation-alternatives></ref><ref id="cit16"><label>16</label><citation-alternatives><mixed-citation xml:lang="ru">Корпус хакасского языка. URL: http://khakas.altaica.ru.</mixed-citation><mixed-citation xml:lang="en">Korpus hakasskogo jazyka. URL: http://khakas.altaica.ru.</mixed-citation></citation-alternatives></ref><ref id="cit17"><label>17</label><citation-alternatives><mixed-citation xml:lang="ru">Корпус якутского языка. URL: http://adictsakha.nsu.ru/corpora/corp.</mixed-citation><mixed-citation xml:lang="en">Korpus jakutskogo jazyka. URL: http://adictsakha.nsu.ru/corpora/corp.</mixed-citation></citation-alternatives></ref><ref id="cit18"><label>18</label><citation-alternatives><mixed-citation xml:lang="ru">Корпус узбекского языка. URL: http://corpus-uz.herokuapp.com.</mixed-citation><mixed-citation xml:lang="en">Korpus uzbekskogo jazyka. URL: http://corpus-uz.herokuapp.com.</mixed-citation></citation-alternatives></ref><ref id="cit19"><label>19</label><citation-alternatives><mixed-citation xml:lang="ru">Корпус шорского и телеутского языков. URL: https://corpora.iea.ras.ru/corpora.</mixed-citation><mixed-citation xml:lang="en">Korpus shorskogo i teleutskogo jazykov. URL: https://corpora.iea.ras.ru/corpora.</mixed-citation></citation-alternatives></ref><ref id="cit20"><label>20</label><citation-alternatives><mixed-citation xml:lang="ru">Лингвистическое ПО «МетаФраз R10». URL: http://www.metafraz.ru.</mixed-citation><mixed-citation xml:lang="en">Lingvisticheskoe PO «MetaFraz R10». URL: http://www.metafraz.ru.</mixed-citation></citation-alternatives></ref><ref id="cit21"><label>21</label><citation-alternatives><mixed-citation xml:lang="ru">C. F. Hockett, Two models of grammatical description, WORD Vol. 10 (1954) 210–234.</mixed-citation><mixed-citation xml:lang="en">C. F. Hockett, Two models of grammatical description, WORD Vol. 10 (1954) 210–234.</mixed-citation></citation-alternatives></ref><ref id="cit22"><label>22</label><citation-alternatives><mixed-citation xml:lang="ru">Yelibayeva G., Sharipbay A., Mukanova A., Razakhova B. (2020) Applied ontology for the automatic classifcation of simple sentences of the Kazakh language. 5th International Conference on Computer Science and Engineering, UBMK 2020. pp. 13-18. doi: 10.1109/UBMK50275.2020.9219461</mixed-citation><mixed-citation xml:lang="en">Yelibayeva G., Sharipbay A., Mukanova A., Razakhova B. (2020) Applied ontology for the automatic classification of simple sentences of the Kazakh language. 5th International Conference on Computer Science and Engineering, UBMK 2020. pp. 13-18. doi: 10.1109/UBMK50275.2020.9219461</mixed-citation></citation-alternatives></ref><ref id="cit23"><label>23</label><citation-alternatives><mixed-citation xml:lang="ru">FrameNet. URL: https://framenet.icsi.berkeley.edu.</mixed-citation><mixed-citation xml:lang="en">FrameNet. URL: https://framenet.icsi.berkeley.edu.</mixed-citation></citation-alternatives></ref><ref id="cit24"><label>24</label><citation-alternatives><mixed-citation xml:lang="ru">Palmer M. (2009). Semlink: Linking PropBank, VerbNet and FrameNet. Proceedings of the Generative Lexicon Conference., pp. 9-15.</mixed-citation><mixed-citation xml:lang="en">Palmer M. (2009). Semlink: Linking PropBank, VerbNet and FrameNet. Proceedings of the Generative Lexicon Conference., pp. 9-15.</mixed-citation></citation-alternatives></ref><ref id="cit25"><label>25</label><citation-alternatives><mixed-citation xml:lang="ru">Gatiatullin A., Suleymanov D., Prokopyev N., Khakimov B. (2020) About turkic morpheme portal. CEUR Workshop Proceedings Institute for history, language and literature, Ufa scientifc center, Russian Academy of Sciences Proceedings of TurkLang 2020, pp. 226-243.</mixed-citation><mixed-citation xml:lang="en">Gatiatullin A., Suleymanov D., Prokopyev N., Khakimov B. (2020) About turkic morpheme portal. CEUR Workshop Proceedings Institute for history, language and literature, Ufa scientific center, Russian Academy of Sciences Proceedings of TurkLang 2020, pp. 226-243.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
