cs.CLOct 1, 2026

BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

Authors: Hasin Almas Sifat

Organizations: Department of Computer Science American International University-Bangladesh Dhaka, Bangladesh

Abstract

Regional linguistic variation remains an important challenge for Bangla natural language processing, particularly in informal and non-standard text. This paper introduces BanglaDial-Abuse, a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla text. The dataset contains 1,000 sentences distributed equally across four linguistic varieties: Standard Bangla, Chattagram, Sylhet, and Barishal, with 250 samples per class. The resource was constructed using a corpus-grounded synthetic procedure incorporating regional variation in pronouns, possessive forms, verb morphology, negation, interrogative structures, postpositions, vocabulary, and Bengali-script spelling conventions while preserving the underlying hostile or abusive meaning. Descriptive analysis shows broadly comparable sentence-length distributions but partially distinct lexical spaces across the four classes. Pairwise Jaccard vocabulary similarity ranges from 0.37 to 0.56. The primary task is four-class regional dialect identification rather than binary abusive-text detection. The dataset is publicly available through Zenodo under a Creative Commons Attribution 4.0 license. The current version is intended as a research and prototyping corpus rather than a native-speaker-validated gold-standard linguistic resource. Keywords: Bangla, Bengali, dialect identification, regional dialect, abusive language, low-resource NLP, Chattagram, Sylhet, Barishal, dataset

Figures & tables

Explore similar work

CardsList
  1. 5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs

    Sep 9, 2026Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed +3DialectsBangla

  2. Unified Multi-Dialectal Neural Machine Translation for Bangla Using the Dwadash Benchmark Corpus

    Aug 12, 2026Rakib Ullah, Md. Ruhul Islam, Tanbir Ahmed +1DialectsNeural Machine Translation