<?xml version='1.0' encoding='utf-8'?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<article article-type="research-article" dtd-version="1.2" xml:lang="ru" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><front><journal-meta><journal-id journal-id-type="issn">2518-1092</journal-id><journal-title-group><journal-title>Research result. Information technologies</journal-title></journal-title-group><issn pub-type="epub">2518-1092</issn></journal-meta><article-meta><article-id pub-id-type="doi">10.18413/2518-1092-2026-11-3-0-3</article-id><article-id pub-id-type="publisher-id">4356</article-id><article-categories><subj-group subj-group-type="heading"><subject>ARTIFICIAL INTELLIGENCE AND DECISION MAKING</subject></subj-group></article-categories><title-group><article-title>&lt;strong&gt;APPLYING MACHINE LEARNING TECHNIQUES&amp;nbsp;TO THE DEVELOPMENT OF ASSISTIVE INTELLIGENT SYSTEM FOR THE HEARING IMPAIRED&lt;/strong&gt;</article-title><trans-title-group xml:lang="en"><trans-title>&lt;strong&gt;APPLYING MACHINE LEARNING TECHNIQUES&amp;nbsp;TO THE DEVELOPMENT OF ASSISTIVE INTELLIGENT SYSTEM FOR THE HEARING IMPAIRED&lt;/strong&gt;</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Velikaya</surname><given-names>Yana Gennadievna</given-names></name><name xml:lang="en"><surname>Velikaya</surname><given-names>Yana Gennadievna</given-names></name></name-alternatives><email>ya.velikaya@apm.ru</email></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Medvedev</surname><given-names>Arsenii Eduardovich</given-names></name><name xml:lang="en"><surname>Medvedev</surname><given-names>Arsenii Eduardovich</given-names></name></name-alternatives><email>1425819@bsuedu.ru</email></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Vassiliev</surname><given-names>Pavel Vladimirovich</given-names></name><name xml:lang="en"><surname>Vassiliev</surname><given-names>Pavel Vladimirovich</given-names></name></name-alternatives><email>geoblock@mail.ru</email></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Mikhelev</surname><given-names>Vladimir Mikhailovich</given-names></name><name xml:lang="en"><surname>Mikhelev</surname><given-names>Vladimir Mikhailovich</given-names></name></name-alternatives><email>mikhelev@bsu.edu.ru</email></contrib></contrib-group><pub-date pub-type="epub"><year>2026</year></pub-date><volume>11</volume><issue>3</issue><fpage>0</fpage><lpage>0</lpage><abstract xml:lang="ru"><p>The paper proposes an intelligent assistive system for Deaf and Hard-of-Hearing users designed to transform acoustic information into a visually accessible form. The study is motivated by the growing prevalence of hearing impairment and by the limitations of conventional rehabilitation tools, which do not fully support the perception of complex acoustic environments. The system is built around a mobile architecture that combines ambient sound analysis, automatic speech recognition, and speech emotion analysis. To map sound to color, the paper introduces an auxiliary textual modality and performs semantic matching between sound descriptions and color descriptions. The experiments show that this approach produces stable sound&amp;ndash;color pairs in HEX format that can be directly integrated into the system. For audio tagging, EfficientAT models were fine-tuned on the AudioSet-Strong dataset; the best performance was achieved with a 2.5-second analysis window and 1.25-second overlap, while the mn10_as model provided the best trade-off between accuracy and computational efficiency. For speech emotion visualization, the continuous Valence-Arousal-Dominance space proved more suitable than rigid categorical labels because it yields a more balanced and less subjective representation of affective states. The results support the use of multimodal learning, transfer learning, and compact neural architectures for practical assistive technologies for users with hearing loss.</p></abstract><trans-abstract xml:lang="en"><p>The paper proposes an intelligent assistive system for Deaf and Hard-of-Hearing users designed to transform acoustic information into a visually accessible form. The study is motivated by the growing prevalence of hearing impairment and by the limitations of conventional rehabilitation tools, which do not fully support the perception of complex acoustic environments. The system is built around a mobile architecture that combines ambient sound analysis, automatic speech recognition, and speech emotion analysis. To map sound to color, the paper introduces an auxiliary textual modality and performs semantic matching between sound descriptions and color descriptions. The experiments show that this approach produces stable sound&amp;ndash;color pairs in HEX format that can be directly integrated into the system. For audio tagging, EfficientAT models were fine-tuned on the AudioSet-Strong dataset; the best performance was achieved with a 2.5-second analysis window and 1.25-second overlap, while the mn10_as model provided the best trade-off between accuracy and computational efficiency. For speech emotion visualization, the continuous Valence-Arousal-Dominance space proved more suitable than rigid categorical labels because it yields a more balanced and less subjective representation of affective states. The results support the use of multimodal learning, transfer learning, and compact neural architectures for practical assistive technologies for users with hearing loss.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>artificial intelligence</kwd><kwd>assistive technologies</kwd><kwd>hearing impairment</kwd><kwd>transfer learning</kwd><kwd>audio tagging</kwd><kwd>automatic speech recognition</kwd><kwd>speech emotion recognition</kwd><kwd>color perception</kwd></kwd-group><kwd-group xml:lang="en"><kwd>artificial intelligence</kwd><kwd>assistive technologies</kwd><kwd>hearing impairment</kwd><kwd>transfer learning</kwd><kwd>audio tagging</kwd><kwd>automatic speech recognition</kwd><kwd>speech emotion recognition</kwd><kwd>color perception</kwd></kwd-group></article-meta></front><back><ref-list><title>Список литературы</title><ref id="B1"><mixed-citation>Allenova O. Why hearing impairment is becoming an epidemic and how to help a hard of hearing child.&amp;nbsp;&amp;ndash; 2025. URL: https://www.kommersant.ru/doc/8137179?ysclid=mun3qme9a243654957 (in Russian).</mixed-citation></ref><ref id="B2"><mixed-citation>Kutsakov A. 2025. GigaAM-v3: open SOTA speech recognition model in Russian. URL: https://habr.com/ru/companies/sberdevices/articles/973160/?ysclid=mun3sqt2c6684552654 (in Russian).</mixed-citation></ref><ref id="B3"><mixed-citation>Medvedev A.E., Mikhelev V.M. Analysis of the principle of cross-modality in artificial intelligence tasks using the CLIP model. Digital, computer and information technologies in science and education. Bryansk. RISO BGU. &amp;ndash; 2. &amp;ndash; 2026. &amp;ndash; 343 p. (in Russian).</mixed-citation></ref><ref id="B4"><mixed-citation>Boltuix. Color-Pedia: a dataset for color naming tasks, palette generation and emotional analysis. Hugging Face. &amp;ndash; 2025. URL: https://huggingface.co/datasets/boltuix/color-pedia</mixed-citation></ref><ref id="B5"><mixed-citation>Chen J., Xiao S., Zhang P., Luo K., Lian D., Liu Z. M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. &amp;ndash; 2025. &amp;ndash; arXiv:2402.03216v5.</mixed-citation></ref><ref id="B6"><mixed-citation>Collopy F. Playing (With) Color. Glimpse. 2(3). &amp;ndash; 2009. &amp;ndash; P. 62-67.</mixed-citation></ref><ref id="B7"><mixed-citation>Fonseca E., Favory X., Pons J., Font F., Serra X. FSD50K: An Open Dataset of Human-Labeled Sound Events. &amp;ndash; 2022. &amp;ndash; ArXiv:2010.00475.</mixed-citation></ref><ref id="B8"><mixed-citation>Gemmeke J.F., Ellis D.P.W., Freedman D., Jansen A., Lawrence W., Moore R. C., Plakal M., Ritter M. Audio Set: An ontology and human-labeled dataset for audio events. Proc. IEEE ICASSP. &amp;ndash; 2017. &amp;ndash; pp. 776-780, doi: 10.1109/ICASSP.2017.7952261.</mixed-citation></ref><ref id="B9"><mixed-citation>Geuder P., Leidinger M. C., von Lupin M., D&amp;ouml;rk M., Schr&amp;ouml;der T. Emosaic: Visualizing Affective Content of Text at Varying Granularity. &amp;ndash; 2020. &amp;ndash; ArXiv:2002.10096.</mixed-citation></ref><ref id="B10"><mixed-citation>Gibson J.J. The senses considered as perceptual systems. Boston: Houghton Mifflin, 1966. &amp;ndash; 335 p.</mixed-citation></ref><ref id="B11"><mixed-citation>Hershey S., Ellis D.P.W., Fonseca E. et al. The Benefit of Temporally-Strong Labels in Audio Event Classification. &amp;ndash; 2021. &amp;ndash; ArXiv:2105.07031.</mixed-citation></ref><ref id="B12"><mixed-citation>Jurafsky D., Martin J.H. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Third Edition. &amp;ndash; 2026. &amp;ndash; 621 р.</mixed-citation></ref><ref id="B13"><mixed-citation>Kondratenko V., Sokolov A., Karpov N., Kutuzov O., Savushkin N., Minkin F. Large Raw Emotional Dataset with Aggregation Mechanism. &amp;ndash; 2022. &amp;ndash; ArXiv:2212.12266.</mixed-citation></ref><ref id="B14"><mixed-citation>Koutini K., Schl&amp;uuml;ter J., Eghbal-zadeh H., Widmer G. Efficient Training of Audio Transformers with Patchout. &amp;ndash; 2022. &amp;ndash; ArXiv:2110.05069.</mixed-citation></ref><ref id="B15"><mixed-citation>Liu Y., Jun E., Q. Li, Heer J. Latent Space Cartography: Visual Analysis of Vector Space Embeddings. Computer Graphics Forum. 38(3). &amp;ndash; 2019. &amp;ndash; 67-78.</mixed-citation></ref><ref id="B16"><mixed-citation>ONNX Community. ONNX: Open Neural Network Exchange. 2026. URL: https://onnx.ai/ (accessed: 26.03.2026).</mixed-citation></ref><ref id="B17"><mixed-citation>Radford A., Kim J. et al. 2021. Learning Transferable Visual Models From Natural Language Supervision. ArXiv:2103.00020.</mixed-citation></ref><ref id="B18"><mixed-citation>Schmid F., Koutini K., Widmer G. Efficient Large-Scale Audio Tagging Via Transformer-to-CNN Knowledge Distillation. 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece, 2023, pp. 1-5, doi: 10.1109/ICASSP49357.2023.10096110</mixed-citation></ref><ref id="B19"><mixed-citation>Schmid F., Koutini K., Widmer G. Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models. &amp;ndash; 2023. &amp;ndash; ArXiv:2310.15648.</mixed-citation></ref><ref id="B20"><mixed-citation>TensorFlow Authors. YAMNet: a model for audio event classification based on MobileNetV1. GitHub. 2026.</mixed-citation></ref><ref id="B21"><mixed-citation>Tzirakis P., Nguyen A., Zafeiriou S., Schuller B.W. Speech Emotion Recognition Using Semantic Information. &amp;ndash; 2021. &amp;ndash; arXiv:2103.02993v1.</mixed-citation></ref><ref id="B22"><mixed-citation>Wagner J., Triantafyllopoulos A., Wierstorf H., Schmitt M., Burkhardt F., Eyben F., Schuller B.W. Dawn of the transformer era in speech emotion recognition: closing the valence gap. IEEE Transactions on Pattern Analysis and Machine Intelligence. &amp;ndash; 2023. &amp;ndash; vol. 45, no. 9, pp. 10745-10759, doi: 10.1109/TPAMI.2023.3263585.</mixed-citation></ref><ref id="B23"><mixed-citation>World Health Organization. World report on hearing. Geneva: WHO, 2021. &amp;ndash; 272 p.</mixed-citation></ref><ref id="B24"><mixed-citation>Wu Y., Chen K., Zhang T. et al. Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation. &amp;ndash; 2024. &amp;ndash; ArXiv:2211.06687.</mixed-citation></ref><ref id="B25"><mixed-citation>Youden W.J. Index for rating diagnostic tests. Cancer. 3(1). &amp;ndash;&amp;nbsp; 1950. &amp;ndash; 32-35.</mixed-citation></ref><ref id="B26"><mixed-citation>Zhang C., Yang Z., He X., Deng L. Multimodal Intelligence: Representation Learning, Information Fusion, and Applications. IEEE Journal of Selected Topics in Signal Processing. 14(3). &amp;ndash; 2020. &amp;ndash; 478-493.</mixed-citation></ref></ref-list></back></article>