OPTIMIZATION OF A HYBRID SEARCH ALGORITHM IN TEXT CORPUSA BASED ON MULTI-LEVEL RELEVANCE METRICS
The relevance of this research stems from the exponential growth of unstructured text data and the increasing demands on the speed and relevance of its processing in such applied areas as news analysis, intelligent search, and natural language processing (NLP). Existing search algorithms often face problems with low accuracy due to the morphological complexity of languages (e.g., Russian), the presence of a large number of noise words, and the need to account for semantic similarity. The goal of this study is to develop and evaluate a hybrid search algorithm that combines lexical match metrics, morphological analysis (stemming), and weighted ranking to improve the completeness and accuracy of searches in text corpora. The scientific innovation lies in the proposed multi-level relevance metric, which integrates exact match counting, word stemming analysis using optimized stemming for the Russian language, and document temporal characteristics. The methodology includes algorithm development, its software implementation, and a comparative analysis of its effectiveness on a real-world dataset, namely, news. The experiment confirmed that the hybrid approach achieves 15-20% higher accuracy than basic methods, especially when working with inflected word forms and compound queries. The study's findings suggest potential for further development of the proposed approach by integrating semantic vector models and deep learning to account for the contextual proximity of concepts.
Levina T.M., Belyakov Ya.V. Optimization of a Hybrid Search Algorithm in Text Corpusa Based on Multi-Level Relevance Metrics // Research result. Information technologies. – Т.11, №3, 2026. – P. 42-48. DOI: 10.18413/2518-1092-2026-11-3-0-4
















While nobody left any comments to this publication.
You can be first.
The references will appear later