<?xml version='1.0' encoding='utf-8'?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<article article-type="research-article" dtd-version="1.2" xml:lang="ru" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><front><journal-meta><journal-id journal-id-type="issn">2518-1092</journal-id><journal-title-group><journal-title>Научный результат. Информационные технологии</journal-title></journal-title-group><issn pub-type="epub">2518-1092</issn></journal-meta><article-meta><article-id pub-id-type="doi">10.18413/2518-1092-2026-11-3-0-3</article-id><article-id pub-id-type="publisher-id">4356</article-id><article-categories><subj-group subj-group-type="heading"><subject>ИСКУССТВЕННЫЙ ИНТЕЛЛЕКТ И ПРИНЯТИЕ РЕШЕНИЙ</subject></subj-group></article-categories><title-group><article-title>&lt;strong&gt;ПРИМЕНЕНИЕ МЕТОДОВ МАШИННОГО ОБУЧЕНИЯ&amp;nbsp;ПРИ СОЗДАНИИ АССИСТИВНОЙ ИНТЕЛЛЕКТУАЛЬНОЙ СИСТЕМЫ ДЛЯ ЛЮДЕЙ С НАРУШЕНИЯМИ СЛУХА&lt;/strong&gt;</article-title><trans-title-group xml:lang="en"><trans-title>&lt;strong&gt;APPLYING MACHINE LEARNING TECHNIQUES&amp;nbsp;TO THE DEVELOPMENT OF ASSISTIVE INTELLIGENT SYSTEM FOR THE HEARING IMPAIRED&lt;/strong&gt;</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Великая</surname><given-names>Яна Геннадьевна</given-names></name><name xml:lang="en"><surname>Velikaya</surname><given-names>Yana Gennadievna</given-names></name></name-alternatives><email>ya.velikaya@apm.ru</email></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Медведев</surname><given-names>Арсений Эдуардович</given-names></name><name xml:lang="en"><surname>Medvedev</surname><given-names>Arsenii Eduardovich</given-names></name></name-alternatives><email>1425819@bsuedu.ru</email></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Васильев</surname><given-names>Павел Владимирович</given-names></name><name xml:lang="en"><surname>Vassiliev</surname><given-names>Pavel Vladimirovich</given-names></name></name-alternatives><email>geoblock@mail.ru</email></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Михелев</surname><given-names>Владимир Михайлович</given-names></name><name xml:lang="en"><surname>Mikhelev</surname><given-names>Vladimir Mikhailovich</given-names></name></name-alternatives><email>mikhelev@bsu.edu.ru</email></contrib></contrib-group><pub-date pub-type="epub"><year>2026</year></pub-date><volume>11</volume><issue>3</issue><fpage>0</fpage><lpage>0</lpage><abstract xml:lang="ru"><p>которые не обеспечивают полноценного восприятия сложной акустической среды. В качестве основы системы рассматривается мобильная архитектура, включающая обработку окружающих звуков, автоматическое распознавание речи и анализ эмоциональной окраски речи. Для сопоставления звука и цвета предложен подход, использующий текстовую промежуточную модальность и семантическое сравнение описаний звуков и цветовых оттенков. Экспериментально показано, что такое сопоставление позволяет получать устойчивые пары &amp;laquo;звук &amp;ndash; цвет&amp;raquo; в HEX-формате, пригодные для дальнейшей интеграции в систему. Для классификации аудио выполнено дообучение моделей семейства EfficientAT на датасете AudioSet-Strong; установлено, что оптимальным является окно анализа 2,5 с при перекрытии 1,25 с, а модель &amp;laquo;mn10_as&amp;raquo; демонстрирует наилучший баланс качества и вычислительной эффективности. Для отображения эмоций речи показано преимущество непрерывного пространства Valence-Arousal-Dominance по сравнению с жесткой категориальной классификацией, поскольку оно дает более равномерное и менее субъективное представление эмоциональных состояний. Полученные результаты подтверждают перспективность применения мультимодальных методов, переноса обучения и компактных нейросетевых моделей для создания практической системы поддержки пользователей с нарушением слуха.</p></abstract><trans-abstract xml:lang="en"><p>The paper proposes an intelligent assistive system for Deaf and Hard-of-Hearing users designed to transform acoustic information into a visually accessible form. The study is motivated by the growing prevalence of hearing impairment and by the limitations of conventional rehabilitation tools, which do not fully support the perception of complex acoustic environments. The system is built around a mobile architecture that combines ambient sound analysis, automatic speech recognition, and speech emotion analysis. To map sound to color, the paper introduces an auxiliary textual modality and performs semantic matching between sound descriptions and color descriptions. The experiments show that this approach produces stable sound&amp;ndash;color pairs in HEX format that can be directly integrated into the system. For audio tagging, EfficientAT models were fine-tuned on the AudioSet-Strong dataset; the best performance was achieved with a 2.5-second analysis window and 1.25-second overlap, while the mn10_as model provided the best trade-off between accuracy and computational efficiency. For speech emotion visualization, the continuous Valence-Arousal-Dominance space proved more suitable than rigid categorical labels because it yields a more balanced and less subjective representation of affective states. The results support the use of multimodal learning, transfer learning, and compact neural architectures for practical assistive technologies for users with hearing loss.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>искусственный интеллект</kwd><kwd>ассистивные технологии</kwd><kwd>проблемы со слухом</kwd><kwd>перенос обучения</kwd><kwd>классификация аудиособытий</kwd><kwd>автоматическое распознавание речи</kwd><kwd>распознавание эмоций речи</kwd><kwd>цветовое восприятие</kwd></kwd-group><kwd-group xml:lang="en"><kwd>artificial intelligence</kwd><kwd>assistive technologies</kwd><kwd>hearing impairment</kwd><kwd>transfer learning</kwd><kwd>audio tagging</kwd><kwd>automatic speech recognition</kwd><kwd>speech emotion recognition</kwd><kwd>color perception</kwd></kwd-group></article-meta></front><back><ref-list><title>Список литературы</title><ref id="B1"><mixed-citation>Алленова О. Почему нарушение слуха становится эпидемией и как помочь слабослышащему ребенку. 18 окт. &amp;ndash; 2025. URL: https://www.kommersant.ru/doc/8137179?ysclid=mun3qme9a243654957</mixed-citation></ref><ref id="B2"><mixed-citation>Куцаков А. 2025. GigaAM-v3: открытая SOTA-модель распознавания речи на русском. 4 дек. URL: https://habr.com/ru/companies/sberdevices/articles/973160/?ysclid=mun3sqt2c6684552654</mixed-citation></ref><ref id="B3"><mixed-citation>Медведев А.Э., Михелёв В.М. Анализ принципа кросс-модальности в задачах искусственного интеллекта на примере модели CLIP. Цифровые, компьютерные и информационные технологии в науке и образовании. 2. &amp;ndash; 2026. &amp;ndash; 343 с.</mixed-citation></ref><ref id="B4"><mixed-citation>Boltuix. Color-Pedia: a dataset for color naming tasks, palette generation and emotional analysis. Hugging Face. &amp;ndash; 2025. URL: https://huggingface.co/datasets/boltuix/color-pedia</mixed-citation></ref><ref id="B5"><mixed-citation>Chen J., Xiao S., Zhang P., Luo K., Lian D., Liu Z. M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. &amp;ndash; 2025. &amp;ndash; arXiv:2402.03216v5.</mixed-citation></ref><ref id="B6"><mixed-citation>Collopy F. Playing (With) Color. Glimpse. 2(3). &amp;ndash; 2009. &amp;ndash; P. 62-67.</mixed-citation></ref><ref id="B7"><mixed-citation>Fonseca E., Favory X., Pons J., Font F., Serra X. FSD50K: An Open Dataset of Human-Labeled Sound Events. &amp;ndash; 2022. &amp;ndash; ArXiv:2010.00475.</mixed-citation></ref><ref id="B8"><mixed-citation>Gemmeke J.F., Ellis D.P.W., Freedman D., Jansen A., Lawrence W., Moore R. C., Plakal M., Ritter M. Audio Set: An ontology and human-labeled dataset for audio events. Proc. IEEE ICASSP. &amp;ndash; 2017. &amp;ndash; pp. 776-780, doi: 10.1109/ICASSP.2017.7952261.</mixed-citation></ref><ref id="B9"><mixed-citation>Geuder P., Leidinger M. C., von Lupin M., D&amp;ouml;rk M., Schr&amp;ouml;der T. Emosaic: Visualizing Affective Content of Text at Varying Granularity. &amp;ndash; 2020. &amp;ndash; ArXiv:2002.10096.</mixed-citation></ref><ref id="B10"><mixed-citation>Gibson J.J. The senses considered as perceptual systems. Boston: Houghton Mifflin, 1966. &amp;ndash; 335 p.</mixed-citation></ref><ref id="B11"><mixed-citation>Hershey S., Ellis D.P.W., Fonseca E. et al. The Benefit of Temporally-Strong Labels in Audio Event Classification. &amp;ndash; 2021. &amp;ndash; ArXiv:2105.07031.</mixed-citation></ref><ref id="B12"><mixed-citation>Jurafsky D., Martin J.H. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Third Edition. &amp;ndash; 2026. &amp;ndash; 621 р.</mixed-citation></ref><ref id="B13"><mixed-citation>Kondratenko V., Sokolov A., Karpov N., Kutuzov O., Savushkin N., Minkin F. Large Raw Emotional Dataset with Aggregation Mechanism. &amp;ndash; 2022. &amp;ndash; ArXiv:2212.12266.</mixed-citation></ref><ref id="B14"><mixed-citation>Koutini K., Schl&amp;uuml;ter J., Eghbal-zadeh H., Widmer G. Efficient Training of Audio Transformers with Patchout. &amp;ndash; 2022. &amp;ndash; ArXiv:2110.05069.</mixed-citation></ref><ref id="B15"><mixed-citation>Liu Y., Jun E., Q. Li, Heer J. Latent Space Cartography: Visual Analysis of Vector Space Embeddings. Computer Graphics Forum. 38(3). &amp;ndash; 2019. &amp;ndash; 67-78.</mixed-citation></ref><ref id="B16"><mixed-citation>ONNX Community. ONNX: Open Neural Network Exchange. 2026. URL: https://onnx.ai/ (дата обращения: 26.03.2026).</mixed-citation></ref><ref id="B17"><mixed-citation>Radford A., Kim J. et al. Learning Transferable Visual Models From Natural Language Supervision. &amp;ndash; 2021. &amp;ndash; ArXiv:2103.00020.</mixed-citation></ref><ref id="B18"><mixed-citation>Schmid F., Koutini K., Widmer G. Efficient Large-Scale Audio Tagging Via Transformer-to-CNN Knowledge Distillation. 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece, 2023, pp. 1-5, doi: 10.1109/ICASSP49357.2023.10096110</mixed-citation></ref><ref id="B19"><mixed-citation>Schmid F., Koutini K., Widmer G. Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models. &amp;ndash; 2023. &amp;ndash; ArXiv:2310.15648.</mixed-citation></ref><ref id="B20"><mixed-citation>TensorFlow Authors. YAMNet: модель для классификации аудиособытий на основе MobileNetV1. GitHub. &amp;ndash; 2026. &amp;ndash;</mixed-citation></ref><ref id="B21"><mixed-citation>Tzirakis P., Nguyen A., Zafeiriou S., Schuller B.W. Speech Emotion Recognition Using Semantic Information. &amp;ndash; 2021. &amp;ndash; arXiv:2103.02993v1.</mixed-citation></ref><ref id="B22"><mixed-citation>Wagner J., Triantafyllopoulos A., Wierstorf H., Schmitt M., Burkhardt F., Eyben F., Schuller B.W. Dawn of the transformer era in speech emotion recognition: closing the valence gap. IEEE Transactions on Pattern Analysis and Machine Intelligence. &amp;ndash; 2023. &amp;ndash; vol. 45, no. 9, pp. 10745-10759, doi: 10.1109/TPAMI.2023.3263585.</mixed-citation></ref><ref id="B23"><mixed-citation>World Health Organization. World report on hearing. Geneva: WHO, 2021. &amp;ndash; 272 p.</mixed-citation></ref><ref id="B24"><mixed-citation>Wu Y., Chen K., Zhang T. et al. Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation. &amp;ndash; 2024. &amp;ndash; ArXiv:2211.06687.</mixed-citation></ref><ref id="B25"><mixed-citation>Youden W.J. Index for rating diagnostic tests. Cancer. 3(1). &amp;ndash;&amp;nbsp; 1950. &amp;ndash; 32-35.</mixed-citation></ref><ref id="B26"><mixed-citation>Zhang C., Yang Z., He X., Deng L. Multimodal Intelligence: Representation Learning, Information Fusion, and Applications. IEEE Journal of Selected Topics in Signal Processing. 14(3). &amp;ndash; 2020. &amp;ndash; 478-493.</mixed-citation></ref></ref-list></back></article>