<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="en"><front><journal-meta><journal-id journal-id-type="publisher-id">Yugra State University Bulletin</journal-id><journal-title-group><journal-title xml:lang="en">Yugra State University Bulletin</journal-title><trans-title-group xml:lang="ru"><trans-title>Вестник Югорского государственного университета</trans-title></trans-title-group></journal-title-group><issn publication-format="print">1816-9228</issn><issn publication-format="electronic">2078-9114</issn><publisher><publisher-name xml:lang="en">Yugra State University</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">10788</article-id><article-id pub-id-type="doi">10.17816/byusu2018037-48</article-id><article-categories><subj-group subj-group-type="toc-heading" xml:lang="en"><subject>Articles</subject></subj-group><subj-group subj-group-type="toc-heading" xml:lang="ru"><subject>Статьи</subject></subj-group><subj-group subj-group-type="article-type"><subject>Research Article</subject></subj-group></article-categories><title-group><article-title xml:lang="en">Information extraction using neural language models for the case of online job listings analysis</article-title><trans-title-group xml:lang="ru"><trans-title>Извлечение информации с использованием нейросетевых моделей языка на примере анализа вакансий в системах онлайн-рекрутмента</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="en"><surname>Botov</surname><given-names>Dmitriy S.</given-names></name><name xml:lang="ru"><surname>Ботов</surname><given-names>Дмитрий Сергеевич</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="en"><p>Senior Lecturer</p></bio><bio xml:lang="ru"><p>Старший преподаватель</p></bio><email>dmbotov@gmail.com</email><xref ref-type="aff" rid="aff1"/></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="en"><surname>Klenin</surname><given-names>Julius D.</given-names></name><name xml:lang="ru"><surname>Кленин</surname><given-names>Юлий Дмитриевич</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="en"><p>Postgraduate student</p></bio><bio xml:lang="ru"><p>Аспирант</p></bio><email>jklen@ya.ru</email><xref ref-type="aff" rid="aff2"/></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="en"><surname>Nikolaev</surname><given-names>Ivan E.</given-names></name><name xml:lang="ru"><surname>Николаев</surname><given-names>Иван Евгеньевич</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="en"><p>Senior Lecturer</p></bio><bio xml:lang="ru"><p>Старший преподаватель</p></bio><email>ivan_nikolaev@csu.ru</email><xref ref-type="aff" rid="aff2"/></contrib></contrib-group><aff-alternatives id="aff1"><aff><institution xml:lang="en">Сhelyabinsk State University</institution></aff><aff><institution xml:lang="ru">Федеральное государственное бюджетное образовательное учреждение высшего образования Челябинский государственный университет</institution></aff></aff-alternatives><aff-alternatives id="aff2"><aff><institution xml:lang="en">Chelyabinsk State University</institution></aff><aff><institution xml:lang="ru">Федеральное государственное бюджетное образовательное учреждение высшего образования Челябинский государственный университет</institution></aff></aff-alternatives><pub-date date-type="pub" iso-8601-date="2018-09-15" publication-format="electronic"><day>15</day><month>09</month><year>2018</year></pub-date><volume>14</volume><issue>3</issue><issue-title xml:lang="en">NO3 (2018)</issue-title><issue-title xml:lang="ru">№3 (2018)</issue-title><fpage>37</fpage><lpage>48</lpage><history><date date-type="received" iso-8601-date="2018-12-27"><day>27</day><month>12</month><year>2018</year></date></history><permissions><copyright-statement xml:lang="en">Copyright ©; 2018, Botov D.S., Klenin J.D., Nikolaev I.E.</copyright-statement><copyright-statement xml:lang="ru">Copyright ©; 2018, Ботов Д.С., Кленин Ю.Д., Николаев И.Е.</copyright-statement><copyright-year>2018</copyright-year><copyright-holder xml:lang="en">Botov D.S., Klenin J.D., Nikolaev I.E.</copyright-holder><copyright-holder xml:lang="ru">Ботов Д.С., Кленин Ю.Д., Николаев И.Е.</copyright-holder><ali:free_to_read xmlns:ali="http://www.niso.org/schemas/ali/1.0/"/><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">http://creativecommons.org/licenses/by-sa/4.0</ali:license_ref></license></permissions><self-uri xlink:href="https://vestnikugrasu.org/byusu/article/view/10788">https://vestnikugrasu.org/byusu/article/view/10788</self-uri><abstract xml:lang="en"><p>In this article we discuss the approach to information extraction (IE) using neural language models. We provide a detailed overview of modern IE methods: both supervised and unsupervised. The proposed method allows to achieve a high quality solution to the problem of analyzing the relevant labor market requirements without the need for a time-consuming labelling procedure. In this experiment, professional standards act as a knowledge base of the labor domain. Comparing the descriptions of work actions and requirements from professional standards with the elements of job listings, we extract four entity types. The approach is based on the classification of vector representations of texts, generated using various neural language models: averaged word2vec, SIF-weighted averaged word2vec, TF-IDF-weighted averaged word2vec, paragraph2vec. Experimentally, the best quality was shown by the averaged word2vec (CBOW) model.</p></abstract><trans-abstract xml:lang="ru"><p>В статье рассматривается подход к извлечению информации с помощью онлайн-обучения на основе определения семантической близости векторов предложений и сущностей базы знаний с помощью нейросетевых моделей языка, обученных без учителя на большом текстовом корпусе предметной области. Приводится подробный обзор современных методов извлечения информации с учителем и без. Предложенный метод позволяет без трудоемкой процедуры разметки текстового корпуса и без применения подходов, основанных на правилах, достичь приемлемого качества в решении задачи анализа актуальных требований рынка труда. В рамках исследования профессиональные стандарты выступают в роли базы знаний предметной области с ограниченной лексикой. В основе подхода лежит определение семантической близости между векторными представлениями текстов, полученных с помощью различных нейросетевых моделей языка: усредненный word2vec, взвешенный по SIF усредненный word2vec, взвешенный по TF-IDF усредненный word2vec, paragraph2vec. В ходе эксперимента лучшее качество работы было показано моделью усредненного word2vec (CBOW).</p></trans-abstract><kwd-group xml:lang="en"><kwd>machine learning</kwd><kwd>natural language processing</kwd><kwd>neural language models</kwd><kwd>classification method</kwd><kwd>information extraction</kwd><kwd>named entity recognition</kwd></kwd-group><kwd-group xml:lang="ru"><kwd>машинное обучение</kwd><kwd>обработка естественного языка</kwd><kwd>нейросетевые модели языка</kwd><kwd>метод классификации</kwd><kwd>извлечение информации</kwd><kwd>распознавание именованных сущностей</kwd></kwd-group><funding-group/></article-meta></front><body></body><back><ref-list><ref id="B1"><label>1.</label><mixed-citation>Shared tasks of the 2015 workshop on noisy user-generated text: Twitter lexical normalization and named entity recognition [Text] / T. Baldwin, de M.-C. Marneffe, B. Han [et al.] // In Proceedings of the Workshop on Noisy User-Generated Text. - Beijing, China, 2015. - P. 126-135.</mixed-citation></ref><ref id="B2"><label>2.</label><mixed-citation>Domain adaptation of rule-based annotators for named-entity recognition tasks [Text] / L. Chiticariu, R. Krishnamurthy, Y. Li [et al.] // Proceedings of the 2010 conference on empirical methods in natural language processing, Association for Computational Linguistics. - San Jose, USA, 2010. - P. 1002-1012.</mixed-citation></ref><ref id="B3"><label>3.</label><mixed-citation>Finkel, J. R. Incorporating non-local information into information extraction systems by gibbs sampling [Text] / J. R. Finkel, T. Grenager, C. Manning // Proceedings of the 43rd annual meeting on association for computational linguistics, Association for Computational Linguistics. - Stanford, USA, 2005. - P. 363-370.</mixed-citation></ref><ref id="B4"><label>4.</label><mixed-citation>Kudo, T. CRF++: Yet another CRF toolkit [Electronic resource] / T. Kudo // GitHub. - URL: https://github.com/taku910/crfpp.</mixed-citation></ref><ref id="B5"><label>5.</label><mixed-citation>Seker, G. A. Extending a CRF-based named entity recognition model for Turkish well formed text and user generated content [Text] / G. A. Seker, G. Eryigit // Semantic Web 8, IOS Press. - 2017. - № 5. - P. 625-642.</mixed-citation></ref><ref id="B6"><label>6.</label><mixed-citation>Bikel, D. M. An algorithm that learns what's in a name [Text] / D. M. Bikel, R. Schwartz, R. M. Weischedel // Machine learning 34. - 1999. - № 1-3. - P. 211-231.</mixed-citation></ref><ref id="B7"><label>7.</label><mixed-citation>Curran, J. R. Language independent NER using a maximum entropy tagger [Text] / J. R. Curran, S. Clark // Proceedings of the seventh conference on Natural language learning at HLT-NAACL. - 2003. - Vol. 4. - P. 164-167.</mixed-citation></ref><ref id="B8"><label>8.</label><mixed-citation>Das, A. Named entity recognition with word embeddings and wikipedia categories for a low-resource language [Text] / A. Das, D. Ganguly, U. Garain // ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP). - USA, New York. - 2017. - Vol. 16, Issue 3. - P. 19-25.</mixed-citation></ref><ref id="B9"><label>9.</label><mixed-citation>Class-based n-gram models of natural language [Text] / P. F. Brown, P. V. Desouza, R. L. Mercer [et al.] // Computational linguistics 18. - 1992. - № 4. - P. 467-479.</mixed-citation></ref><ref id="B10"><label>10.</label><mixed-citation>Siencnik, S. K. Adapting word2vec to named entity recognition [Text] / S. K. Siencnik // Proceedings of the 20th Nordic Conference of Computational Linguistics. - Sweden, 2015. - № 109. - P. 239-243.</mixed-citation></ref><ref id="B11"><label>11.</label><mixed-citation>Wu, Y. A study of neural word embeddings for named entity recognition in clinical text [Text] / Y. Wu, J. Xu, M. Jiang, Y. Zhang, H. Xu // AMIA Annual Symposium Proceedings, American Medical Informatics Association. - USA, San Francisco. - 2015. - Vol. 2015. - P. 1326-1333.</mixed-citation></ref><ref id="B12"><label>12.</label><mixed-citation>Toral, A. A proposal to automatically build and maintain gazetteers for Named Entity Recognition by using Wikipedia [Text] / A. Toral, R. Munoz // Proceedings of the Workshop on NEW TEXT Wikis and blogs and other dynamic text sources. - Italy, Trento. - 2006. - Vol. 1. - P. 56-61.</mixed-citation></ref><ref id="B13"><label>13.</label><mixed-citation>Chiu, J. P.-C. Named entity recognition with bidirectional LSTM-CNNs [Text] / J. P.-C. Chiu, E. Nichols // Transactions of the Association for Computational Linguistics. - 2016. - Vol. 4. - P. 357-370.</mixed-citation></ref><ref id="B14"><label>14.</label><mixed-citation>Huang, Z. Bidirectional LSTM-CRF models for sequence tagging [Text] / Z. Huang, W. Xu, K. Yu // arXiv preprint arXiv:1508.01991. - 2015.</mixed-citation></ref><ref id="B15"><label>15.</label><mixed-citation>RELigator: chemical-disease relation extraction using prior knowledge and textual information [Text] / E. Pons, B. F. H. Becker, S. A. Akhondi [et al.] // Proceedings of the Fifth BioCreative Challenge Evaluation Workshop. - Spain, Sevilla. - 2015. - Vol. 1. - P. 247-253.</mixed-citation></ref><ref id="B16"><label>16.</label><mixed-citation>Relation classification via convolutional deep neural network [Text] / D. Zeng, K. Liu, S. Lai [et al.] // Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers. - Ireland, Dublin. - 2014. - Vol. 1. - P. 2335-2344.</mixed-citation></ref><ref id="B17"><label>17.</label><mixed-citation>Plank, B. Embedding semantic similarity in tree kernels for domain adaptation of relation extraction [Text] / B. Plank, A. Moschitti // Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics. Volume 1: Long Papers. - Bulgaria, Sofia. - 2013. - Vol. 1. - P. 1498-1507.</mixed-citation></ref><ref id="B18"><label>18.</label><mixed-citation>Quan, C. An unsupervised text mining method for relation extraction from biomedical literature [Text] / C. Quan, M. Wang, F. Ren // PloS one - 2014. - Vol. 9, Issue 7. - P. 1-8.</mixed-citation></ref><ref id="B19"><label>19.</label><mixed-citation>Extracting Relational Facts by an End-to-End Neural Model with Copy Mechanism [Text] / X. Zeng, D. Zeng, S. He [et al.] // Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. - Australia, Melbourne. - 2018. - Vol. 1. - P. 506-514.</mixed-citation></ref><ref id="B20"><label>20.</label><mixed-citation>Xu, B. CN-DBpedia: A never-ending Chinese knowledge extraction system [Text] / B. Xu, Y. Xu, J. Liang [et al.] // In International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, France, Arras : Springer International Publishing. - 2017. - Vol. 1, Part II, LNAI 10351 - P. 428-438.</mixed-citation></ref><ref id="B21"><label>21.</label><mixed-citation>Exploring Encoder-Decoder Model for Distant Supervised Relation Extraction [Text] / S. Su, N. Jia, X. Cheng [et al.] // Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI-18). - Sweden, Stockholm. - 2018. - Vol. 1. - P. 4389-4395.</mixed-citation></ref><ref id="B22"><label>22.</label><mixed-citation>Distributed representations of words and phrases and their compositionality [Text] / T. Mikolov, I. Sutskever, K. Chen [et al.] // In Advances in neural information processing systems. - 2013. - Vol. 1. - P. 3111-3119.</mixed-citation></ref><ref id="B23"><label>23.</label><mixed-citation>Le, Q. Distributed representations of sentences and documents [Text] / Q. Le, T. Mikolov // International Conference on Machine Learning. - China, Beijing. - 2014. - Vol. 32. - P.1188-1196.</mixed-citation></ref><ref id="B24"><label>24.</label><mixed-citation>Arora, S. A simple but tough-to-beat baseline for sentence embeddings [Text] / S. Arora, Y. Liang, T. Ma // International Conference on Learning Representations (ICLR). - France, Toulon. - 2017. - Vol. 1. - P. 1-16.</mixed-citation></ref><ref id="B25"><label>25.</label><mixed-citation>Rehurek, R. Software framework for topic modelling with large corpora [Text] / R. Rehurek, P. Sojka // Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks. - Malta, Valletta. - 2010. - Vol. 1. - P. 46-50.</mixed-citation></ref><ref id="B26"><label>26.</label><mixed-citation>RUSSE’2018: a Shared Task on Word Sense Induction for the Russian Language [Text] / A. Panchenko, A. Lopukhina, D. Ustalov [et al.] // Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference «Dialogue 2018». - Russia, Moscow. - 2018. - Vol. 1. - P. 547-564.</mixed-citation></ref><ref id="B27"><label>27.</label><mixed-citation>RUSSE: The First Workshop on Russian Semantic Similarity [Text] / A. Panchenko, N. V. Loukachevitch, D. Ustalov [et al.] // Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference «Dialogue 2015». - Russia, Moscow. - 2015. - Vol. 2. - P. 89-105.</mixed-citation></ref><ref id="B28"><label>28.</label><mixed-citation>FactRuEval 2016: Evaluation of Named Entity Recognition and Fact Extraction Systems for Russian [Text] / V. V. Bocharov, S. V. Alexeeva, A. A. Bodrova [et al.] // Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference «Dialogue 2016». - Russia, Moscow. - 2016. - Vol. 1. - P. 702-720.</mixed-citation></ref><ref id="B29"><label>29.</label><mixed-citation>Пархоменко, П. А. Обзор и экспериментальное сравнение методов кластеризации текстов [Текст] / П. А. Пархоменко, А. А. Григорьев, Н. А. Астраханцев // Труды ИСП РАН. - 2017. - Т. 29, Вып. 2. - С. 161-200.</mixed-citation></ref></ref-list></back></article>
