<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="en"><front><journal-meta><journal-id journal-id-type="publisher-id">Yugra State University Bulletin</journal-id><journal-title-group><journal-title xml:lang="en">Yugra State University Bulletin</journal-title><trans-title-group xml:lang="ru"><trans-title>Вестник Югорского государственного университета</trans-title></trans-title-group></journal-title-group><issn publication-format="print">1816-9228</issn><issn publication-format="electronic">2078-9114</issn><publisher><publisher-name xml:lang="en">Yugra State University</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">705229</article-id><article-id pub-id-type="doi">10.18822/byusu202602110-117</article-id><article-categories><subj-group subj-group-type="toc-heading" xml:lang="en"><subject>Mathematical modeling and information technology</subject></subj-group><subj-group subj-group-type="toc-heading" xml:lang="ru"><subject>Математическое моделирование и информационные технологии</subject></subj-group><subj-group subj-group-type="article-type"><subject>Research Article</subject></subj-group></article-categories><title-group><article-title xml:lang="en">Adaptive semantically-oriented quantization of deep neural networks</article-title><trans-title-group xml:lang="ru"><trans-title>Адаптивное семантико-ориентированное квантование глубоких нейронных сетей</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="en"><surname>Shaptsev</surname><given-names>Valeriy A.</given-names></name><name xml:lang="ru"><surname>Шапцев</surname><given-names>Валерий Алексеевич</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="en"><p>Doctor of Engineering Science, Professor of the Information Systems Department</p></bio><bio xml:lang="ru"><p>доктор технических наук, профессор Школы компьютерных наук</p></bio><email>vashaptsev@ya.ru</email><xref ref-type="aff" rid="aff1"/></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="en"><surname>Khrupin</surname><given-names>Danila S.</given-names></name><name xml:lang="ru"><surname>Хрупин</surname><given-names>Данила Станиславович</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="en"><p>Postgraduate student of the Information Systems Department</p></bio><bio xml:lang="ru"><p>аспирант Школы компьютерных наук</p></bio><email>Khrupin24@mail.ru</email><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff-alternatives id="aff1"><aff><institution xml:lang="en">Tyumen State University</institution></aff><aff><institution xml:lang="ru">Тюменский государственный университет</institution></aff></aff-alternatives><pub-date date-type="pub" iso-8601-date="2026-06-30" publication-format="electronic"><day>30</day><month>06</month><year>2026</year></pub-date><volume>22</volume><issue>2</issue><issue-title xml:lang="en"/><issue-title xml:lang="ru"/><fpage>110</fpage><lpage>117</lpage><history><date date-type="received" iso-8601-date="2026-03-31"><day>31</day><month>03</month><year>2026</year></date><date date-type="accepted" iso-8601-date="2026-05-10"><day>10</day><month>05</month><year>2026</year></date></history><permissions><copyright-statement xml:lang="en">Copyright ©; 2026, Yugra State University</copyright-statement><copyright-statement xml:lang="ru">Copyright ©; 2026, Югорский государственный университет</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="en">Yugra State University</copyright-holder><copyright-holder xml:lang="ru">Югорский государственный университет</copyright-holder><ali:free_to_read xmlns:ali="http://www.niso.org/schemas/ali/1.0/"/><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by-sa/4.0</ali:license_ref></license></permissions><self-uri xlink:href="https://vestnikugrasu.org/byusu/article/view/705229">https://vestnikugrasu.org/byusu/article/view/705229</self-uri><abstract xml:lang="en"><p>Subject of research: an algorithm for adaptive semantics-oriented quantization of deep neural networks that ensures the separability of objects with a high degree of visual similarity.</p> <p>Purpose of research: to develop an adaptive semantics-oriented quantization algorithm with minimal loss of recognition accuracy for objects in hard-to-distinguish classes and a high degree of compression in the digital implementation.</p> <p>Research methods: numerical analysis of the normalized matrix of identification errors for objects of overlapping classes. Evaluation of the contribution of weight tensors to class separability in feature space to determine the semantic significance of neural network layers. Heterogeneous distribution of the bit depth of weight coefficients and intermediate activations based on estimates of the cumulative sensitivity curve parameters. Quantization-Aware Training procedure with selective freezing of detection layers.</p> <p>Objects of research: methods for compressing deep neural networks in computer vision.</p> <p>Research findings: a heterogeneous bit-depth distribution, preserving the FP32 (32-bit floating-point) format only for 29 % of semantically significant layers, allows maintaining recognition accuracy (Mean Average Precision, mAP@0.5) at 98.6 % of the accuracy of the baseline, unquantized model. The method proposed in this work, tested on the YOLOv8n architecture, achieves a 2.78-fold compression of the FP32 model. Furthermore, for hard-to-distinguish objects, it outperforms the standard PTQ (Post-Training Quantization) and the homogeneous approach QAT (Quantization-Aware Training) in terms of recall (the proportion of correctly detected objects of a given class) and the F1 score (the harmonic mean of precision and recall, in the range [0, 1]).</p></abstract><trans-abstract xml:lang="ru"><p>Предмет исследования: алгоритм адаптивного семантико-ориентированного квантования глубоких нейросетей, обеспечивающий разделимость объектов, имеющих высокую степень визуального сходства.</p> <p>Цель исследования: разработка алгоритма адаптивного семантико-ориентированного квантования с минимальной потерей точности распознавания объектов трудноразличимых классов и высокой степенью компрессии цифровой реализации.</p> <p>Методы исследования: численный анализ нормализованной матрицы ошибок идентификации объектов пересекающихся классов. Оценка вклада тензоров весов в разделимость классов в пространстве признаков для определения семантической значимости слоёв нейросети. Гетерогенное распределение разрядности весовых коэффициентов и промежуточных активаций по оценкам параметров кривой кумулятивной чувствительности. Процедура обучения Quantization-Aware Training с селективной фиксацией слоёв детекции.</p> <p>Объекты исследования: способы компрессии глубоких нейронных сетей компьютерного зрения.</p> <p>Основные результаты исследования: гетерогенное распределение разрядности с сохранением формата FP32 (Floating Point 32 бита с плавающей запятой) лишь для 29 % семантически значимых слоёв позволяет удержать точность распознавания (mean Average Precision, mAP@0,5) на уровне 98,6 % от точности базовой, неквантованной модели. Предложенный в работе метод, апробированный на архитектуре YOLOv8n, обеспечивает сжатие модели FP32 в 2,78 раза. При этом на трудноразличимых объектах имеет место превосходство над стандартным подходом PTQ (Post-Training Quantization) и гомогенным QAT (Quantization-Aware Training, обучение с квантованием) по полноте (Recall – доля верно обнаруженных объектов данного класса) и F1-мере (среднее гармоническое точности и полноты в интервале [0,1]).</p></trans-abstract><kwd-group xml:lang="en"><kwd>neural network quantization</kwd><kwd>mixed precision</kwd><kwd>YOLOv8</kwd><kwd>QAT</kwd><kwd>PTQ</kwd><kwd>semantic significance</kwd><kwd>mathematical formalization</kwd><kwd>hard-to-distinguish objects</kwd></kwd-group><kwd-group xml:lang="ru"><kwd>квантование нейронных сетей</kwd><kwd>смешанная точность</kwd><kwd>YOLOv8</kwd><kwd>QAT</kwd><kwd>PTQ</kwd><kwd>семантическая значимость</kwd><kwd>математическая формализация</kwd><kwd>трудноразличимые объекты</kwd></kwd-group><funding-group/></article-meta></front><body></body><back><ref-list><ref id="B1"><label>1.</label><mixed-citation>Хрупин, Д. С. Метод квантования нейронных сетей обнаружения на встраиваемых системах / Д. С. Хрупин, В. А. Шапцев // Научный результат. Информационные технологии. – 2025. – Т. 10, № 4. – С. 72–78.</mixed-citation></ref><ref id="B2"><label>2.</label><mixed-citation>A survey of quantization methods for efficient neural network inference / A. Gholami, S. kim, Z. Dong [et al.] // Low-power computer vision. – London : Chapman and Hall/CRC, 2022. – P. 291–326.</mixed-citation></ref><ref id="B3"><label>3.</label><mixed-citation>Bochkovskiy, A. YOLOv4: Optimal Speed and Accuracy of Object Detection / A. Bochkovskiy, C. Y. Wang, H. Y. M. Liao // arXiv. – URL: https://arxiv.org/abs/2004.10934 (date of access: 20.03.2026).</mixed-citation></ref><ref id="B4"><label>4.</label><mixed-citation>BRECQ: Pushing the limit of post-training quantization by block reconstruction / Y. Li, R. Gong, X. Tan [et al.] // arXiv. – URL: https://arxiv.org/abs/2102.05426 (date of access: 20.03.2026).</mixed-citation></ref><ref id="B5"><label>5.</label><mixed-citation>Data-free quantization through weight equalization and bias correction / M. Nagel, M. van Baalen, T. Blankevoort, M. Welling // Proceedings of the IEEE/CVF international conference on computer vision. – Seoul, Korea (South), 2019. – P. 1325–1334.</mixed-citation></ref><ref id="B6"><label>6.</label><mixed-citation>Distance-IoU loss: Faster and better learning for bounding box regression / Z. Zheng, P. Wang, W. Liu [et al.] // Proceedings of the AAAI conference on artificial intelligence. – 2020. – Vol. 34, № 7. – P. 12993–13000.</mixed-citation></ref><ref id="B7"><label>7.</label><mixed-citation>Explore Ultralytics YOLOv8 – Ultralytics YOLO Docs // ultralytics. – URL: https://docs.ultralytics.com/models/yolov8/ (date of access: 24.02.2026).</mixed-citation></ref><ref id="B8"><label>8.</label><mixed-citation>Fisher-aware Quantization for DETR Detectors with Critical-category Objectives / H. Yang, Y. Huang, Z. Dong [et al.] // arXiv. – URL: https://arxiv.org/abs/2407.03442 (date of access: 20.03.2026).</mixed-citation></ref><ref id="B9"><label>9.</label><mixed-citation>Focal loss for dense object detection / T. Y. Lin, P. Goyal, R. Girshik [et al.] // Proceedings of the IEEE international conference on computer vision. – Venice, Italy, 2017. – P. 2980–2988.</mixed-citation></ref><ref id="B10"><label>10.</label><mixed-citation>Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection / X. Li, W. Wang, L. Wu [et al.] // Advances in neural information processing systems. – 2020. – Vol. 33. – P. 21002–21012.</mixed-citation></ref><ref id="B11"><label>11.</label><mixed-citation>HAWQ: Hessian aware quantization of neural networks with mixed-precision / Z. Dong, Z. Yao, A. Gholami [et al.] // Proceedings of the IEEE/CVF international conference on computer vision. – Seoul, Korea (South), 2019. – С. 293-302.</mixed-citation></ref><ref id="B12"><label>12.</label><mixed-citation>HAWQ-V2: Hessian aware trace-weighted quantization of neural networks / Z. Dong, Z. Yao, Y. Cai [et al.] // Advances in neural information processing systems. – 2020. – Vol. 33. – P. 18518–18529.</mixed-citation></ref><ref id="B13"><label>13.</label><mixed-citation>MIO-TCD: A new benchmark dataset for vehicle classification and localization / Z. Luo, C. Lemarie, S. Li [et al.] // IEEE Transactions on Image Processing. – 2018. – Vol. 27, № 10. – P. 5129–5141.</mixed-citation></ref><ref id="B14"><label>14.</label><mixed-citation>PyTorch: An imperative style, high-performance deep learning library / A. Paszke, S. Gross, F. Massa [et al.] // Advances in neural information processing systems. – 2019. – Vol. 32. – P. 1–12.</mixed-citation></ref><ref id="B15"><label>15.</label><mixed-citation>QReg: On Regularization Effects of Quantization / M. H. AskariHemmat, R. A. Hemmat, A. Hoffman [et al.] // arXiv. – URL: https://arxiv.org/abs/2206.12372 (date of access: 20.03.2026).</mixed-citation></ref><ref id="B16"><label>16.</label><mixed-citation>Quantization and training of neural networks for efficient integer-arithmetic-only inference / B. Jacob, S. Kligys, B. Chen [et al.] // Proceedings of the IEEE conference on computer vision and pattern recognition. – Salt Lake City, 2018. – P. 2704–2713.</mixed-citation></ref><ref id="B17"><label>17.</label><mixed-citation>Reg-ptq: Regression-specialized post-training quantization for fully quantized object detector / Ding Y. et al. // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. – Seattle, 2024. – P. 16174–16184.</mixed-citation></ref><ref id="B18"><label>18.</label><mixed-citation>The pascal visual object classes (voc) challenge / M. Everingham, L. Van Gool, Ch. K. I. Williams [et al.] // International journal of computer vision. – 2010. – Vol. 88, № 2. – P. 303–338.</mixed-citation></ref><ref id="B19"><label>19.</label><mixed-citation>ZeroQ: A novel zero shot quantization framework / Y. Cai,Z. Yao, Z. Dong [et al.] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. – Seattle, 2020. – P. 13169–13178.</mixed-citation></ref></ref-list></back></article>
