Paper ID: 2305.10848

Advancing Full-Text Search Lemmatization Techniques with Paradigm Retrieval from OpenCorpora

Dmitriy Kalugin-Balashov

In this paper, we unveil a groundbreaking method to amplify full-text search lemmatization, utilizing the OpenCorpora dataset and a bespoke paradigm retrieval algorithm. Our primary aim is to streamline the extraction of a word's primary form or lemma - a crucial factor in full-text search. Additionally, we propose a compact dictionary storage strategy, significantly boosting the speed and precision of lemma retrieval.

Submitted: May 18, 2023