Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving long-form prosody and consistent single-speaker narration uncovered. We present BanglaKontho, a single-speaker Bangla TTS corpus of 20 hours derived from professional audiobook recordings: 7,050 segmented utterances with verified transcripts at 24 kHz. We also release a reusable Bangla text normalizer covering Bangladeshi-style digit grouping, currency and date expressions, Danda punctuation and Unicode normalization, together with the full preprocessing pipeline. An MB-iSTFT-VITS baseline trained from scratch reaches 9.5% WER and 4.46 naturalness MOS, against 16.0% and 3.16 for the same architecture retrained on the 12-hour IndicTTS-Bn corpus. The corpus is released openly under CC BY-NC 4.0.
Figures & tables
Corpus
Hours
Speakers
Domain
SR (kHz)
Access
OpenSLR Bengali [ 1 ]
∼ 200
505
Read prompts
16
Open
IndicTTS-Bn [ 5 ]
12
1
Read prompts (studio)
48
Open
Common Voice Bn [ 2 ]
60+
many
Read prompts
48
Open
Shrutilipi (Bn portion) [ 7 ]
437
many
Broadcast
16
Open
BanglaKontho \xspace (this work)
20
1
Audiobook
24
Open (CC BY-NC 4.0)
Table 1: Bangla speech corpora most commonly used as TTS data sources. BanglaKontho \xspace fills the long-form, single-speaker audiobook slot. The IndicTTS-Bn row describes the single 12-hour Bangla voice used as our baseline, not the full Bangla release.
Property
Value
Total duration
\SI 20 \hour
Utterances
7,050
Speakers
1
Source domain
Audiobook
Sampling rate (release)
\SI 24000 \hertz
Sampling rate (baseline)
\SI 22050 \hertz
Table 2: BanglaKontho \xspace dataset statistics. The chapter-disjoint split (no chapter contributes utterances to more than one of train / val / test) prevents same-chapter prosodic leakage between training and evaluation.
System
MCD ↓
WER ↓
MOS ↑
BanglaKontho \xspace (ours)
5.85
9.5
4.46±0.06
BanglaKontho \xspace w/o norm.
5.81
14.0
3.80±0.08
IndicTTS-Bn
5.97
16.0
3.16±0.09
Natural speech
–
8.0
4.90±0.03
Table 3: Comparison of our VITS-based model trained on BanglaKontho \xspace against the same architecture trained on IndicTTS-Bn ( \SI 12 \hour ) and against natural speech. MCD is in dB and WER in %, both lower-is-better; WER is computed with IndicWav2Vec on the synthesized audio. Naturalness MOS is averaged over 15 native Bangla raters following the ITU-T P.800 protocol [ 23 ] and reported with a 95% confidence interval.