Bosh sahifa

AUTOMATIC DETECTION OF CULTURAL LACUNAE USING CORPUS TOOLS: A CORPUS-DRIVEN PIPELINE FOR ENGLISH-UZBEK COMPARABLE CORPORA

ID: GEN-2026-316DOI 10.5281/zenodo.20284280CC-BY-4.0

Authors (1)

Received

Received

Revised

Revised

Accepted

Accepted

Published

May 19, 2026

Abstract

Cultural lacunae are lexical or conceptual gaps where a culture-bound term has no direct equivalent in another language. Traditional identification methods are subjective and non-scalable. This study proposes a corpus-driven pipeline for automatic detection of cultural lacunae using corpus tools and cross-lingual embeddings. Two comparable corpora (American English and Uzbek, 5 million words each) were constructed. The pipeline detected 147 candidate lacunae with strict precision of 72% and lenient precision of 88.5%. Food, social rituals, and legal-administrative domains showed the highest lacuna density. Building on Ataboev s (2019a, 2019b, 2020, 2024a, 2024b) corpus linguistics research, this study extends automatic detection to cultural gap identification.

Keywords

Original

cultural lacunaecorpus linguisticsautomatic detectioncomparable corporaUzbek corpus

Cite this article

qizi, R.M.M. (2026). AUTOMATIC DETECTION OF CULTURAL LACUNAE USING CORPUS TOOLS: A CORPUS-DRIVEN PIPELINE FOR ENGLISH-UZBEK COMPARABLE CORPORA. Research and Publications. https://doi.org/10.5281/zenodo.20284280

References

  1. [1]Ataboev, N. B. (2019a). ICT in linguistic studies: Application of electronic language corpus and corpus-based analysis. Test Engineering and Management, *81*, 4170-4176.
  2. [2]Ataboev, N. B. (2019b). Problematic issues of corpus analysis and its shortcomings. ISJ Theoretical & Applied Science, *10*(78), 170-173.
  3. [3]Ataboev, N. B. (2020). Corpus-based research on the language features of corpus linguistics: In the example of ECOCL. Language, *3*(2139), 950.
  4. [4]Ataboev, N. (2024a). Analysis of the media texts corpus in the prism of existing English diachronic corpora [Media matnlar korpusining mavjud ingliz tili diaxron korpuslari prizmasida tahlili]. Acta NUUz, *1*(1.2.1), 288
  5. [5]https://doi.org/10.69617/nuuz.v1i1.2.1.1242
  6. [6]Ataboev, N. (2024b). Diachronic corpora: The role of corpus linguistics methodology in the studies on language development [Diaxronik korpuslar: Til rivoji tadqiqida korpus lingvistikasi metodologiyasining o‘rni]. Acta NUUz, *1*(1.3), 272
  7. [7]https://doi.org/10.69617/nuuz.v1i1.3.1385
  8. [8]Ataboev, N. (2024c). Media texts as the main resource of language social expression and language enrichment. Foreign Languages in Uzbekistan, *2024*(1), 60-75.