BanglaBox: A Phonetically-Balanced Corpus and Data-Efficient Foundation-Model Adaptation for Bangla Text-to-Speech with Zero-Shot Voice Cloning
Abstract
We present a recipe for adapting English-pretrained autoregressive TTS foundation models to underrepresented languages, demonstrated on Bangladeshi Bangla. Existing Bangla TTS corpora are small and single-speaker, and to our knowledge no open zero-shot voice-cloning system is available for the Bangladeshi register. We contribute a phonetically- and gender-balanced two-tier Bangladeshi Bangla corpus balanced via a tiered Jensen-Shannon divergence objective over conjunct clusters (juktakkhor), together with three fine-tuning changes: a merge-consistent tokenizer extension, Bangla text normalization, and a prompt-masked dual-loss objective. These changes preserve zero-shot cloning across the language switch. Our BanglaEval protocol applies Wilcoxon signed-rank tests with Bonferroni correction over native-speaker ratings. BanglaBox attains near-natural Naturalness, outperforms commercial and open-source baselines on Naturalness, Speaker Similarity, and Clarity, and reaches speaker similarity comparable to prior few-shot results while using approximately 7x less Bangla fine-tuning audio. We further validate the recipe beyond our own test split using a seven-category stress set of difficult real-world text, the public BnTTS evaluation benchmarks, and naturally occurring Bangla that no language model wrote. All artifacts, including the corpus, weights, code, and complete evaluation materials, are released publicly and unconditionally.
Figures & tables
| Partition | Spk. | Hours | Utts. |
|---|---|---|---|
| By tier | |||
| Tier A (studio) | 13 | 367.3 | 220,000 |
| Tier B (curated) | 23 | 153.0 | 92,000 |
| By split (speaker-disjoint test) | |||
| Train (11A + 20B) | 31 | 422 | 253,000 |
| Dev (within-spk) | 31 | 22 | 13,000 |
| Dataset | Yr | Hrs | Spk | MS | Cl | Dom. |
| SUBESCO | 2021 | 7.7 | 20 | ✓ | ✗ | Emot. |
| SHRUTI | 2011 | 22 | 34 | ✓ | ✗ | Read |
| SUST TTS | 2021 | 30 | 1 | ✗ | ✗ | Read |
| CommonVoice (v9) | 2022 | 56 | – | ✓ | ✗ | Crowd |
| Widely used non-English TTS corpora | ||||||
| LJSpeech | 2017 | 24 | 1 | ✗ | ✗ | Audiobook |
| System | Nat. | SMOS | Clar. |
|---|---|---|---|
| Google TTS | 4.27 0.17 | 4.17 0.13 | 4.51 0.13 |
| Azure TTS | 4.13 0.14 | 4.07 0.11 | 4.44 0.15 |
| Narakeet | 3.97 0.16 | 3.71 0.13 | 4.02 0.12 |
| IndicF5 | 3.89 0.15 | 3.55 0.14 | 3.95 0.12 |
| YourTTS (FT) | 4.18 0.14 | 3.08 0.15 | 4.39 0.13 |
| XTTS-v2 (FT) | 4.21 | — | 4.35 |
| System | NISQA | CER | SECS |
|---|---|---|---|
| Google TTS | 4.52 | 2.2 | — |
| Azure TTS | 4.43 | 2.7 | — |
| Narakeet | 4.37 | 3.3 | — |
| IndicF5 | 4.13 | 3.8 | 0.37 |
| YourTTS (FT) | 3.69 | 3.1 | 0.44 |
| XTTS-v2 (FT) | 4.02 | 3.6 | 0.47 |
| System | Ref. | SECS | SMOS |
| IndicF5 | 5 s | 0.31 | 3.46 0.14 |
| 10 s | 0.37 | 3.55 0.14 | |
| XTTS-v2 (FT) | 5 s | 0.41 | 3.88 0.15 |
| 10 s | 0.47 | 3.79 0.13 | |
| YourTTS (FT) | 5 s | 0.42 | 3.01 0.15 |
| 10 s | 0.44 | 3.08 0.15 |
| Configuration | Nat. | CER | SECS |
|---|---|---|---|
| Full BanglaBox | 4.56 0.16 | 2.9 | 0.627 |
| Tokenizer ext. | 4.39 ∗∗ 0.17 | 3.3 | 0.60 ∗ |
| Text normalization | 4.23 ∗∗∗ 0.17 | 4.4 | 0.56 ∗∗ |
| Prompt mask | 4.34 ∗∗ 0.16 | 3.9 | 0.51 ∗∗ |
| Trainable enc. | 4.55 0.15 | 3.0 | 0.62 |
| 100 h subset | 4.18 0.18 | 4.3 | 0.54 |
| h subset | Conj. | Rare | Nat. | CER | SECS |
|---|---|---|---|---|---|
| JSD-balanced | 97.0% | 91.7% | 4.18 0.18 | 4.3 | 0.54 |
| Random | 93.5% | 79.8% | 4.21 0.13 | 7.1 | 0.52 |
| JSD-balanced, 3 seeds | — | — | 4.31 | — | — |
| Random, 3 seeds | — | — | 4.18 | — | — |
| Category | CER (%) | SECS |
|---|---|---|
| Out-of-vocabulary words | 2.2 | .72 |
| Spelling errors | 2.4 | .72 |
| Proper names | 1.9 | .72 |
| Code-mixed English | 1.3 | .71 |
| Complex juktakkhor | 0.7 | .73 |
| Numerics | 1.0 | .72 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
|---|---|
| Base model | Chatterbox TTS (T3) |
| Architecture | Llama-style; 1024-dim, 30 layers, 16 heads |
| Trainable / frozen | T3 (520M) / voice enc. + S3Gen |
| Optimizer | AdamW ( , ) |
| Peak LR | , cosine, 5% warmup |
| Weight decay |
| Step | Output |
|---|---|
| 0. input | \banglafont সাঃ ৩:৪ এর মধ্যে 12 টি Apple খেতে চাই। |
| 1. digit conv. | \banglafont সাঃ ৩:৪ এর মধ্যে ১২ টি Apple খেতে চাই। |
| 2. colon-num | \banglafont সাঃ ৩ এর ৪ এর মধ্যে ১২ টি Apple খেতে চাই। |
| 3. num word | \banglafont সাঃ তিন এর চার এর মধ্যে বারো টি Apple খেতে চাই। |
| 4. Unicode norm | (no change) |
| 5. honorific | \banglafont সাল্লাল্লাহু আলাইহি ওয়া সাল্লাম তিন এর চার এর মধ্যে বারো টি Apple খেতে চাই। |
| Benchmark | BnTTS CER | Our CER | Our SECS |
|---|---|---|---|
| BengaliStimuli53 | 8.6% | 7.2% | .78 |
| BengaliNamedEntity1000 | 4.0% | 2.9% | .68 |
| ShortText | 9.2% | 6.7% | .51 |
| Evaluation text | CER (%) |
|---|---|
| LLM-written | 2.2 |
| Natural Tier B | 2.5 |
| Domain | Nat. | SMOS | Clar. |
|---|---|---|---|
| News | 4.47 | 4.54 | 4.69 |
| Customer care | 4.56 | 4.34 | 4.69 |
| Teaching | 4.58 | 4.55 | 4.69 |
| Healthcare | 4.74 | 4.50 | 4.75 |
| E-commerce | 4.41 | 4.55 | 4.70 |
| Finance | 4.59 | 4.50 | 4.59 |
| Nat. | SMOS | Clar. | ||||
|---|---|---|---|---|---|---|
| BanglaBox vs. | ||||||
| 1.6e-4 | 0.50 | 8.9e-8 | 0.72 | 1.1e-5 | 0.60 | |
| Azure | 1.9e-8 | 0.76 | 1.4e-11 | 0.90 | 1.6e-4 | 0.51 |
| Narakeet | 3.1e-9 | 0.79 | 1.9e-13 | 0.99 | 1.7e-13 | 0.98 |
| IndicF5 | 7.3e-11 | 0.87 | 3.8e-14 | 1.00 | 9.7e-14 | 0.99 |
| YourTTS | 3.5e-7 | 0.67 | 3.9e-14 | 1.00 | 1.2e-6 | 0.65 |
| System | Model / voice / settings |
|---|---|
| Gemini-TTS gemini-2.5-flash-tts (bn-BD), | |
| voice Achernar , 22.05 kHz, rate 1.0 | |
| Azure | bn-BD-NabanitaNeural |
| Narakeet | voice Mahiya (Bangladeshi accent) |
| Conjunct group | Phoneme acc. (%) | Linguist-confirmed errors |
|---|---|---|
| Frequent ( 5%) | 97.7 | 23 |
| Moderate (1–5%) | 94.3 | 35 |
| Rare ( 0.3%) | 88.2 | 29 |
| Requirement / metric | Chunked | Single-pass |
|---|---|---|
| Premature EOS — truncation | 3.3% | 14.7% |
| Quality — CER, cap | 2.9% | 6.4% |
| Quality — CER, over-cap | 5.3% | 21.4% |
| Content drift — CER by length | 3.4% ‡ | 3.4% 29% |
| Voice drift — start–end SECS | 0.75 | 0.83 † |
| Efficiency — RTF (H100) | 1.7 | 0.66 |
| Condition | F0 RMSE (Hz) | Dur. dev. (%) |
|---|---|---|
| BanglaBox — short utts. | 13.2 | 2.4 |
| BanglaBox — long utts. | 17.1 | 4.8 |
| Gate | Rejected | Rate |
|---|---|---|
| 1. length | 461 | 0.22% |
| 2. TTR | 83 | 0.04% |
| 3. dedup | 873 | 0.42% |
| 4. lang-ID | 21 | 0.01% |
| 5. perplexity | 317 | 0.15% |
| 6. phonetic novelty | n/a | n/a |
| Domain | Hours | Domain | Hours |
|---|---|---|---|
| News | 62.4 | E-commerce | 47.8 |
| Customer care | 58.8 | Finance | 47.7 |
| Teaching | 55.1 | IT | 47.8 |
| Healthcare | 47.7 | ||
| Total (Tier A) = 367.3 hour (studio-recorded) | |||
| Gate | Threshold / criterion | Why (what it blocks) |
|---|---|---|
| 1. length | 6 – 28 words (token count) | too short/long for TTS |
| 2. TTR | type-token ratio on -word strings | repetitive text |
| 3. dedup | cosine sim against existing corpus | near-duplicates |
| 4. lang-ID | Bengali probability (fastText 176-lang ID) | wrong-language text |
| 5. perplexity | Bangla LM perplexity | unnatural/broken Bangla |
| 6. phonetic novelty | gain on the current under-covered juktakkhor set | adds no new coverage |
| Weights | Max cov. | Scripts |
|---|---|---|
| change | changed | |
| paper | — | — |
| unif. | 0.1 pp | 1.13% |
| swap | 0.1 pp | 0.29% |
| swap | 0.1 pp | 0.84% |
| reversed | 0.1 pp | 1.53% |
| Dataset | License | Spk. (ret./init.) | Hours (ret./init.) |
|---|---|---|---|
| SUBAK.KO | CC BY 4.0 | 7/11 | 54.30/87.34 |
| BanSpeech | CC BY 4.0 | 1/1 | 5.14/6.52 |
| MediBeng | CC BY 4.0 | 1/2 | 3.67/7.11 |
| IndicVoices (Bengali) | CC BY 4.0 | 1/2 | 5.95/20.00 |
| Shrutilipi (Bengali) | CC BY 4.0 | 1/3 | 2.70/18.37 |
| Common Voice Bengali | CC0 1.0 | 3/8 | 3.00/15.00 |
| Abbrev. | Full expansion (Bangla) |
|---|---|
| \banglafont সাঃ | \banglafont সাল্লাল্লাহু আলাইহি ওয়া সাল্লাম |
| \banglafont রাঃ | \banglafont রাদিআল্লাহু আনহু |
| \banglafont রহঃ | \banglafont রাহিমাহুল্লাহ |
| \banglafont হাঃ | \banglafont হাফিজাহুল্লাহ |
| \banglafont মাঃ | \banglafont মাদ্দাজিল্লুহু |
| \banglafont দাঃ বাঃ | \banglafont দামাত বারাকাতুহুম |
| Term | Definition |
|---|---|
| ARR | ACL Rolling Review submission cycle. |
| AR | Autoregressive (transformer decoder). |
| BPE | Byte-Pair Encoding (subword tokenization). |
| CER | Character Error Rate (ASR transcription error). |
| CI | 95% Confidence Interval. |
| ECAPA-TDNN | Emphasized Channel Attention, Propagation and Aggregation TDNN. |