<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="en"><front><journal-meta><journal-id journal-id-type="publisher-id">Yugra State University Bulletin</journal-id><journal-title-group><journal-title xml:lang="en">Yugra State University Bulletin</journal-title><trans-title-group xml:lang="ru"><trans-title>Вестник Югорского государственного университета</trans-title></trans-title-group></journal-title-group><issn publication-format="print">1816-9228</issn><issn publication-format="electronic">2078-9114</issn><publisher><publisher-name xml:lang="en">Yugra State University</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">706463</article-id><article-id pub-id-type="doi">10.18822/byusu20260271-76</article-id><article-categories><subj-group subj-group-type="toc-heading" xml:lang="en"><subject>Mathematical modeling and information technology</subject></subj-group><subj-group subj-group-type="toc-heading" xml:lang="ru"><subject>Математическое моделирование и информационные технологии</subject></subj-group><subj-group subj-group-type="article-type"><subject>Research Article</subject></subj-group></article-categories><title-group><article-title xml:lang="en">Development and research of a data preprocessing system LM2-REC for recommendation systems</article-title><trans-title-group xml:lang="ru"><trans-title>Разработка и исследование системы предобработки данных LM2-REC для рекомендательных систем</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="en"><surname>Zamyshlyaeva</surname><given-names>Alyona A.</given-names></name><name xml:lang="ru"><surname>Замышляева</surname><given-names>Алёна Александровна</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="en"><p>Doctor of Physics and Mathematics, Professor, Director of the Institute of Natural Sciences and Mathematics</p></bio><bio xml:lang="ru"><p>доктор физико-математических наук, профессор, директор Института естественных и точных наук</p></bio><email>zamyshliaevaaa@susu.ru</email><xref ref-type="aff" rid="aff1"/></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="en"><surname>Kim</surname><given-names>Artem L.</given-names></name><name xml:lang="ru"><surname>Ким</surname><given-names>Артём Леонидович</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="en"><p>Postgraduate student at the Institute of Natural and Exact Sciences</p></bio><bio xml:lang="ru"><p>аспирант Института естественных и точных наук</p></bio><email>ar.kim@mail.ru</email><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff-alternatives id="aff1"><aff><institution xml:lang="en">South Ural State University (National Research University)</institution></aff><aff><institution xml:lang="ru">Южно-Уральский государственный университет (Национальный исследовательский университет)</institution></aff></aff-alternatives><pub-date date-type="pub" iso-8601-date="2026-06-30" publication-format="electronic"><day>30</day><month>06</month><year>2026</year></pub-date><volume>22</volume><issue>2</issue><issue-title xml:lang="en"/><issue-title xml:lang="ru"/><fpage>71</fpage><lpage>76</lpage><history><date date-type="received" iso-8601-date="2026-04-20"><day>20</day><month>04</month><year>2026</year></date><date date-type="accepted" iso-8601-date="2026-05-09"><day>09</day><month>05</month><year>2026</year></date></history><permissions><copyright-statement xml:lang="en">Copyright ©; 2026, Yugra State University</copyright-statement><copyright-statement xml:lang="ru">Copyright ©; 2026, Югорский государственный университет</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="en">Yugra State University</copyright-holder><copyright-holder xml:lang="ru">Югорский государственный университет</copyright-holder><ali:free_to_read xmlns:ali="http://www.niso.org/schemas/ali/1.0/"/><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by-sa/4.0</ali:license_ref></license></permissions><self-uri xlink:href="https://vestnikugrasu.org/byusu/article/view/706463">https://vestnikugrasu.org/byusu/article/view/706463</self-uri><abstract xml:lang="en"><p>Subject of research: methods for preprocessing textual data in recommender systems based on the integration of large language models (LLMs) and recurrent neural networks.</p> <p>Purpose of research: to develop and validate the LM2-Rec data preprocessing method that extracts an expanded set of semantic and categorical features from user-generated text.</p> <p>Research methods: semantic analysis using a locally deployed LLM, text vectorization via a MacBERT encoder, training of an LSTM recurrent neural network for categorical feature prediction, comparative analysis against baseline methods (TF-IDF with logistic regression, embeddings with logistic regression and random forest), and a series of few-shot learning experiments.</p> <p>Objects of research: the LM2-Rec data preprocessing system comprising a pipeline of a local LLM, a MacBERT encoder, and an LSTM network; the Amazon Reviews Dataset of textual user reviews.</p> <p>Research findings: the LM2-Rec method achieves an F1-Score of 0.9170 with a processing time of 2.87 s, which is five times faster than the embeddings-with-logistic-regression baseline. In few-shot learning scenarios (50–500 examples), the method maintains consistently high accuracy (F1-Score above 0.95), confirming its robustness under limited-data conditions. The LLM-based extraction of semantic features yields a completeness rate exceeding 90 % across all key fields.</p></abstract><trans-abstract xml:lang="ru"><p>Предмет исследования: методы предобработки текстовых данных для рекомендательных систем на основе комбинации больших языковых моделей (LLM) и нейронной сети LSTM.</p> <p>Цель исследования: разработка и проверка эффективности метода предобработки текстовых данных LM2-Rec, обеспечивающего извлечение широкого набора семантических и категориальных признаков.</p> <p>Методы исследования: семантический анализ с применением локальной LLM, векторизация текста с использованием энкодера paraphrase-multilingual-MiniLM-L12-v2, обучение нейронной сети LSTM для предсказания категориальных признаков, сравнительный анализ с базовыми методами (TF-IDF, эмбеддинги с логистической регрессией и случайным лесом), а также эксперименты в few-shot сценарии.</p> <p>Объекты исследования: система предобработки данных LM2-Rec, включающая комбинацию обработок на основе локальной LLM, энкодера paraphrase-multilingual-MiniLM-L12-v2 и LSTM-сети; набор текстовых данных Amazon Review Dataset.</p> <p>Основные результаты исследования: разработан метод LM2-Rec, обеспечивающий F1-Score 0,9170 при скорости обработки 2.87 с, что в 5 раз быстрее подхода «эмбеддинги + логистическая регрессия». В сценарии few-shot обучения (50–500 примеров) метод показывает стабильно высокую точность (F1-Score 0,95+), подтверждая эффективность в условиях малых данных. Полнота LLM-извлечения семантических признаков составляет свыше 90 % по всем ключевым полям.</p></trans-abstract><kwd-group xml:lang="en"><kwd>computer science</kwd><kwd>data preprocessing</kwd><kwd>neural networks</kwd><kwd>LSTM</kwd><kwd>LLM</kwd><kwd>recommender systems</kwd></kwd-group><kwd-group xml:lang="ru"><kwd>компьютерные науки</kwd><kwd>предобработка данных</kwd><kwd>нейронные сети</kwd><kwd>LSTM</kwd><kwd>LLM</kwd><kwd>рекомендательные системы</kwd></kwd-group><funding-group/></article-meta></front><body></body><back><ref-list><ref id="B1"><label>1.</label><mixed-citation>Attention is all you need / A. Vaswani, N. Shazeer, N. Parmar [et al.] // Advances in Neural Information Processing Systems. – 2017. – Vol. 30. – P. 5998–6008.</mixed-citation></ref><ref id="B2"><label>2.</label><mixed-citation>Bengio, Y. Representation learning: A review and new perspectives / Y. Bengio, A. Courville, P. Vincent // IEEE Transactions on Pattern Analysis and Machine Intelligence. – 2013. – Vol. 35, № 8. – P. 1798–1828.</mixed-citation></ref><ref id="B3"><label>3.</label><mixed-citation>Biau, G. A random forest guided tour / G. Biau, E. Scornet // TEST. – 2016. – Vol. 25, № 2. – P. 197–227.</mixed-citation></ref><ref id="B4"><label>4.</label><mixed-citation>Efficient estimation of word representations in vector space / T. Mikolov, K. Chen, G. Corrado, J. Dean // Proceedings of the International Conference on Learning Representations (ICLR 2013) (Scottsdale, Arizona, USA, May 2–4, 2013). – Scottsdale, 2013. – P. 1–12.</mixed-citation></ref><ref id="B5"><label>5.</label><mixed-citation>Generalizing from a few examples: A survey on few-shot learning / Y. Wang, Q. Yao, J. T. Kwok, L. M. Ni // ACM Computing Surveys. – 2020. – Vol. 53, № 3. – P. 1–34.</mixed-citation></ref><ref id="B6"><label>6.</label><mixed-citation>Han, J. Data mining: concepts and techniques / J. Han, M. Kamber, J. Pei. – 3rd ed. – Waltham : Morgan Kaufmann, 2012. – 703 p.</mixed-citation></ref><ref id="B7"><label>7.</label><mixed-citation>He, R. Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering / R. He, J. McAuley // Proceedings of the 25th International Conference on World Wide Web (WWW’16) (Montreal, Canada, April 11–15, 2016). – New York : ACM, 2016. – P. 507–517.</mixed-citation></ref><ref id="B8"><label>8.</label><mixed-citation>Hochreiter, S. Long short-term memory / S. Hochreiter, J. Schmidhuber // Neural Computation. – 1997. – Vol. 9, № 8. – P. 1735–1780.</mixed-citation></ref><ref id="B9"><label>9.</label><mixed-citation>Krizhevsky, A. ImageNet classification with deep convolutional neural networks / A. Krizhevsky, I. Sutskever, G. E. Hinton // Advances in Neural Information Processing Systems. – 2012. – Vol. 25. – P. 1097–1105.</mixed-citation></ref><ref id="B10"><label>10.</label><mixed-citation>Language models are few-shot learners / T. B. Brown, B. Mann, N. Ryder [et al.] // Advances in Neural Information Processing Systems. – 2020. – Vol. 33. – P. 1877–1901.</mixed-citation></ref><ref id="B11"><label>11.</label><mixed-citation>LeCun, Y. Deep learning / Y. LeCun, Y. Bengio, G. Hinton // Nature. – 2015. – Vol. 521. – P. 436–444.</mixed-citation></ref><ref id="B12"><label>12.</label><mixed-citation>Llama: Open and efficient foundation language models / H. Touvron, T. Lavril, G. Izacard [et al.] // arXiv preprint. – 2023. – URL: https://arxiv.org/abs/2302.13971 (date of access: 15.12.2025).</mixed-citation></ref><ref id="B13"><label>13.</label><mixed-citation>Neural collaborative filtering / X. He, L. Liao, H. Zhang [et al.] // Proceedings of the 26th International Conference on World Wide Web (WWW’17) (Perth, Australia, April 3–7, 2017). – New York : ACM, 2017. – P. 173–182.</mixed-citation></ref><ref id="B14"><label>14.</label><mixed-citation>Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing / P. Liu, W. Yuan, J. Fu [et al.] // ACM Computing Surveys. – 2023. – Vol. 55, № 9. – P. 1–35.</mixed-citation></ref><ref id="B15"><label>15.</label><mixed-citation>Qaiser, S. Text mining: use of TF-IDF to examine the relevance of words to documents / S. Qaiser, R. Ali // International Journal of Computer Applications. – 2018. – Vol. 181, № 1. – P. 25–29.</mixed-citation></ref><ref id="B16"><label>16.</label><mixed-citation>Sherstinsky, A. Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network / A. Sherstinsky // Physica D: Nonlinear Phenomena. – 2020. – Vol. 404. – P. 132306.</mixed-citation></ref><ref id="B17"><label>17.</label><mixed-citation>Sokolova, M. A systematic analysis of performance measures for classification tasks / M. Sokolova, G. Lapalme // Information Processing and Management. – 2009. – Vol. 45, № 4. – P. 427–437.</mixed-citation></ref></ref-list></back></article>
