We present a recipe for adapting English-pretrained autoregressive TTS foundation models to underrepresented languages, demonstrated on Bangladeshi Bangla. Existing Bangla TTS corpora are small and single-speaker, and to our knowledge no open zero-shot voice-cloning system is available for the Bangladeshi register. We contribute a phonetically- and gender-balanced two-tier Bangladeshi Bangla corpus balanced via a tiered Jensen-Shannon divergence objective over conjunct clusters (juktakkhor), together with three fine-tuning changes: a merge-consistent tokenizer extension, Bangla text normalization, and a prompt-masked dual-loss objective. These changes preserve zero-shot cloning across the language switch. Our BanglaEval protocol applies Wilcoxon signed-rank tests with Bonferroni correction over native-speaker ratings. BanglaBox attains near-natural Naturalness, outperforms commercial and open-source baselines on Naturalness, Speaker Similarity, and Clarity, and reaches speaker similarity comparable to prior few-shot results while using approximately 7x less Bangla fine-tuning audio. We further validate the recipe beyond our own test split using a seven-category stress set of difficult real-world text, the public BnTTS evaluation benchmarks, and naturally occurring Bangla that no language model wrote. All artifacts, including the corpus, weights, code, and complete evaluation materials, are released publicly and unconditionally.
Figures & tables
Figure 1: Two-tier corpus construction (§ 2 ). Tier A passes 9 quality gates (thresholds: Table 21 , App. C.3 ) and two-pass human review before studio recording; Tier B is filtered by signal-to-noise ratio (SNR) and an ablation check (16/39 speakers dropped). Tiers merge under the tiered Jensen–Shannon divergence (JSD) objective (Eq. 1 ).
Partition
Spk.
Hours
Utts.
By tier
Tier A (studio)
13
367.3
∼ 220,000
Tier B (curated)
23
153.0
∼ 92,000
By split (speaker-disjoint test)
Train (11A + 20B)
31
∼ 422
∼ 253,000
Dev (within-spk)
31
∼ 22
∼ 13,000
Table 1: Corpus statistics by tier and split ( dˉ=6.0 s mean utterance duration).
Dataset
Yr
Hrs
Spk
MS
Cl
Dom.
SUBESCO
2021
7.7
20
✓
✗
Emot.
SHRUTI
2011
22
34
✓
✗
Read
SUST TTS
2021
30
1
✗
✗
Read
CommonVoice (v9)
2022
56
–
✓
✗
Crowd
Widely used non-English TTS corpora
LJSpeech
2017
24
1
✗
✗
Audiobook
Table 2: Comparison with existing Bangla speech datasets (MS = multi-speaker; Cl = cloning). BanglaBox is ∼9× Common Voice Bengali v9 and ∼17× the SUST TTS Corpus. Widely used non-English corpora are included for scale; none applies phonetic balancing, whereas BanglaBox balances phone, diphone, triphone and juktakkhor distributions.
Figure 2: Tiered JSD across coverage-feedback iterations 1, 5, 10, 14; smaller polygon ⇒ closer to target T .
Figure 3: Corpus analysis. (a) Utterance durations peak at 4–6 s (green). (b) Word counts peak at 7–8 words (orange). (c) Vocabulary follows sub-linear Heaps’ law growth. Histograms in (a) and (b) cover the ∼ 241 k utterances of duration ≥ 2 s; bottom pills summarize statistics for the full 520.3 h / ∼ 312 k-utterance corpus.
Figure 4: BanglaBox architecture (§ 3.1 ). Only T3 is trainable ( ∇ ); the voice encoder and S3Gen vocoder stay frozen ( ∗ ). The first 75 prompt tokens are masked from Lspeechmasked . Inset: 704 pretrained embedding rows copied, 3,536 new rows mean-initialized from centroid μ .
Figure 5: BanglaEval evaluation pipeline (§ 4 ). ∗ XTTS-v2 (FT) is scored on objective and cloning metrics only. Character error rate (CER); speaker-encoder cosine similarity (SECS); speaker-similarity mean opinion score (SMOS). Results in Table 3 .
System
Nat. ↑
SMOS ↑
Clar. ↑
Google TTS
4.27 ± 0.17
4.17 ± 0.13
4.51 ± 0.13
Azure TTS
4.13 ± 0.14
4.07 ± 0.11
4.44 ± 0.15
Narakeet
3.97 ± 0.16
3.71 ± 0.13
4.02 ± 0.12
IndicF5
3.89 ± 0.15
3.55 ± 0.14
3.95 ± 0.12
YourTTS (FT)
4.18 ± 0.14
3.08 ± 0.15
4.39 ± 0.13
XTTS-v2 (FT)
4.21
—
4.35
Table 3: Subjective results from 24 raters (76 utts., 1,824 ratings/system). Bold : best synthesized; † human upper bound. ∗∗p<0.01 vs. baselines (Wilcoxon + Bonferroni). CI computation, Krippendorff’s α and exact statistics: App. 14 . The XTTS-v2 (FT) row is new in this revision.
System
NISQA ↑
CER ↓
SECS ↑
Google TTS
4.52
2.2
—
Azure TTS
4.43
2.7
—
Narakeet
4.37
3.3
—
IndicF5
4.13
3.8
0.37
YourTTS (FT)
3.69
3.1
0.44
XTTS-v2 (FT)
4.02
3.6
0.47
Table 4: Objective results: NISQA Mittag et al. (2021) (reference-free MOS), CER (Whisper large-v3, Bengali), SECS (ECAPA-TDNN, 192-dim) over 436 utts. SECS n/a for commercial (no speaker adaptation). Markers: ∗∗p<0.01 vs. each cloning-capable baseline (Wilcoxon signed-rank on per-utterance SECS, Bonferroni over three comparisons; App. 14 ).
System
Ref.
SECS ↑
SMOS ↑
IndicF5
5 s
0.31
3.46 ± 0.14
10 s
0.37
3.55 ± 0.14
XTTS-v2 (FT)
5 s
0.41
3.88 ± 0.15
10 s
0.47
3.79 ± 0.13
YourTTS (FT)
5 s
0.42
3.01 ± 0.15
10 s
0.44
3.08 ± 0.15
Table 5: Zero-shot voice cloning on held-out speakers, 5 s vs. 10 s reference. Bold : best overall. Markers: ∗p<0.05 , ∗∗p<0.01 vs. each cloning-capable baseline (Wilcoxon + Bonferroni). Metric defs in § 4 .
Figure 6: Data scaling on 100/200/520 h subsets. Naturalness (purple, left) keeps improving; character error rate (CER, orange, right) saturates beyond 200 h. Stars: full 520 h configuration.
Configuration
Nat. ↑
CER ↓
SECS ↑
Full BanglaBox
4.56 ± 0.16
2.9
0.627
− Tokenizer ext.
4.39 ∗∗ ± 0.17
3.3
0.60 ∗
− Text normalization
4.23 ∗∗∗ ± 0.17
4.4
0.56 ∗∗
− Prompt mask
4.34 ∗∗ ± 0.16
3.9
0.51 ∗∗
+ Trainable enc.
4.55 ± 0.15
3.0
0.62
100 h subset
4.18 ± 0.18
4.3
0.54
Table 6: Ablation study. Upper: each row removes one component; “ + Trainable enc.” un-freezes the voice encoder and “ − Tokenizer ext.” reverts to a naive extension. Lower: data-scaling subsets. Bold : full BanglaBox. Markers: ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 (Wilcoxon + Bonferroni).
100 h subset
Conj. ↑
Rare ↑
Nat. ↑
CER ↓
SECS ↑
JSD-balanced
97.0%
91.7%
4.18 ± 0.18
4.3
0.54
Random
93.5%
79.8%
4.21 ± 0.13
7.1
0.52
JSD-balanced, 3 seeds
—
—
4.31
—
—
Random, 3 seeds
—
—
4.18
—
—
Table 7: Balanced vs. random selection at matched 100 h from the same Tier A pool, trained identically. Coverage: share of the conjunct inventory covered, in full and over its rare tail (App. B.7 ). Seed rows (1234, 2025, 31337) are new in this revision.
Category
CER (%) ↓
SECS ↑
Out-of-vocabulary words
2.2
.72
Spelling errors
2.4
.72
Proper names
1.9
.72
Code-mixed English
1.3
.71
Complex juktakkhor
0.7
.73
Numerics
1.0
.72
Table 8: Stress test over seven categories of difficult real-world text.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Base model
Chatterbox TTS (T3)
Architecture
Llama-style; 1024-dim, 30 layers, 16 heads
Trainable / frozen
T3 (520M) / voice enc. + S3Gen
Optimizer
AdamW ( β1=0.9 , β2=0.999 )
Peak LR
5×10−6 , cosine, 5% warmup
Weight decay
1×10−2
Appendix
Table 9: Fine-tuning and inference configuration for BanglaBox.
Step
Output
0. input
\banglafont সাঃ ৩:৪ এর মধ্যে 12 টি Apple খেতে চাই।
1. digit conv.
\banglafont সাঃ ৩:৪ এর মধ্যে ১২ টি Apple খেতে চাই।
2. colon-num
\banglafont সাঃ ৩ এর ৪ এর মধ্যে ১২ টি Apple খেতে চাই।
3. num → word
\banglafont সাঃ তিন এর চার এর মধ্যে বারো টি Apple খেতে চাই।
4. Unicode norm
(no change)
5. honorific
\banglafont সাল্লাল্লাহু আলাইহি ওয়া সাল্লাম তিন এর চার এর মধ্যে বারো টি Apple খেতে চাই।
Appendix
Table 10: Worked example tracing one input through the 6-step normalization pipeline of § 3.1 Stage 2.
Benchmark
BnTTS CER
Our CER
Our SECS
BengaliStimuli53
8.6%
7.2%
.78
BengaliNamedEntity1000
4.0%
2.9%
.68
ShortText
9.2%
6.7%
.51
Appendix
Table 11: External benchmark comparison against reported BnTTS figures.
Evaluation text
CER (%) ↓
LLM-written
2.2
Natural Tier B
2.5
Appendix
Table 12: Evaluation-text register: model-written versus naturally occurring Bangla.
Domain
Nat. ↑
SMOS ↑
Clar. ↑
News
4.47
4.54
4.69
Customer care
4.56
4.34
4.69
Teaching
4.58
4.55
4.69
Healthcare
4.74
4.50
4.75
E-commerce
4.41
4.55
4.70
Finance
4.59
4.50
4.59
Appendix
Table 13: BanglaBox per-domain subjective scores (mean over 24 raters, 10–11 items per domain). Naturalness stays within 4.41 – 4.74 across all 7 domains.
Nat.
SMOS
Clar.
BanglaBox vs.
p
r
p
r
p
r
Google
1.6e-4
0.50
8.9e-8
0.72
1.1e-5
0.60
Azure
1.9e-8
0.76
1.4e-11
0.90
1.6e-4
0.51
Narakeet
3.1e-9
0.79
1.9e-13
0.99
1.7e-13
0.98
IndicF5
7.3e-11
0.87
3.8e-14
1.00
9.7e-14
0.99
YourTTS
3.5e-7
0.67
3.9e-14
1.00
1.2e-6
0.65
Appendix
Table 14: Exact two-sided Wilcoxon p -values and rank-biserial effect sizes r on per-item mean differences ( n=76 ). Bonferroni threshold α/5=0.01 ; every BanglaBox-vs.-baseline comparison clears it on all 3 scales. † Human vs. BanglaBox.
System
Model / voice / settings
Google
Gemini-TTS gemini-2.5-flash-tts (bn-BD),
voice Achernar , 22.05 kHz, rate 1.0
Azure
bn-BD-NabanitaNeural
Narakeet
voice Mahiya (Bangladeshi accent)
Appendix
Table 15: Commercial baseline documentation: model versions, voices, and settings. All commercial audio was generated in a single archived batch, queried between Jan 13 and Feb 7, 2026.
Conjunct group
Phoneme acc. (%) ↑
Linguist-confirmed errors
Frequent ( > 5%)
97.7
23
Moderate (1–5%)
94.3
35
Rare ( < 0.3%)
88.2
29
Appendix
Table 16: Juktakkhor pronunciation accuracy by training-frequency group, over the 436 -utterance objective set. Accuracy is phoneme-level; the error column counts only cases a native linguist confirmed as genuine mispronunciations after forced alignment flagged them.
Requirement / metric
Chunked
Single-pass
Premature EOS — truncation
3.3%
14.7%
Quality — CER, ≤ cap
2.9%
6.4%
Quality — CER, over-cap
5.3%
21.4%
Content drift — CER by length
3.4% ‡
3.4% → 29%
Voice drift — start–end SECS
0.75
0.83 †
Efficiency — RTF (H100)
1.7
0.66
Appendix
Table 17: Chunked vs. single-pass long-form synthesis over 40 multi-sentence paragraphs spanning both sides of the 850 -token cap. † Single-pass often stops early, so its start–end SECS covers incomplete audio and overstates consistency. ‡ All items.
Condition
F0 RMSE (Hz) ↓
Dur. dev. (%) ↓
BanglaBox — short utts.
13.2
2.4
BanglaBox — long utts.
17.1
4.8
Appendix
Table 18: Prosody vs. held-out ground truth ( 56 utterances with both a real recording and a BanglaBox synthesis; n=28 per row). Duration deviation is the mean absolute difference from reference.
Gate
Rejected
Rate
1. length
461
0.22%
2. TTR
83
0.04%
3. dedup
873
0.42%
4. lang-ID
21
0.01%
5. perplexity
317
0.15%
6. phonetic novelty
n/a
n/a
Appendix
Table 19: Per-gate rejection: 6,807 of 206,943 candidates ( 3.29% ). Gates 6 and 9 are prompt-level steering, not per-candidate predicates.
Domain
Hours
Domain
Hours
News
62.4
E-commerce
47.8
Customer care
58.8
Finance
47.7
Teaching
55.1
IT
47.8
Healthcare
47.7
Total (Tier A) = 367.3 hour (studio-recorded)
Appendix
Table 20: Tier A 7-domain text distribution.
Gate
Threshold / criterion
Why (what it blocks)
1. length
6 – 28 words (token count)
too short/long for TTS
2. TTR
type-token ratio ≥0.55 on ≥12 -word strings
repetitive text
3. dedup
cosine sim ≤0.85 against existing corpus
near-duplicates
4. lang-ID
Bengali probability ≥0.93 (fastText 176-lang ID)
wrong-language text
5. perplexity
Bangla LM perplexity ∈[20,240]
unnatural/broken Bangla
6. phonetic novelty
gain ≥+5% on the current under-covered juktakkhor set
adds no new coverage
Appendix
Table 21: The 9 automated quality gates: threshold and the failure each one blocks. Gate 6 (phonetic novelty) and gate 9 (gap-unit) operationalise the coverage-feedback loop.
Weights
Max cov.
Scripts
change
changed
(.1,.2,.3,.4) paper
—
—
(.25×4) unif.
< 0.1 pp
1.13%
(.1,.2,.4,.3) swap t/j
< 0.1 pp
0.29%
(.2,.1,.3,.4) swap p/d
< 0.1 pp
0.84%
(.4,.3,.2,.1) reversed
< 0.1 pp
1.53%
Appendix
Table 22: Weight sensitivity of the tiered-JSD selection loop, measured against the paper’s setting. Coverage change is the largest shift across the four tiers; scripts changed is the share of admitted scripts that differ. Coverage is unchanged on every tier under every perturbation.
Dataset
License
Spk. (ret./init.)
Hours (ret./init.)
SUBAK.KO
CC BY 4.0
7/11
54.30/87.34
BanSpeech
CC BY 4.0
1/1
5.14/6.52
MediBeng
CC BY 4.0
1/2
3.67/7.11
IndicVoices (Bengali)
CC BY 4.0
1/2
5.95/20.00
Shrutilipi (Bengali)
CC BY 4.0
1/3
2.70/18.37
Common Voice Bengali
CC0 1.0
3/8
3.00/15.00
Appendix
Table 23: Tier B sources with licenses and per-dataset retention (speakers and hours, retained/initial) after the SNR and ablation-driven filter passes. Speaker counts denote distinct voices after listening-based merging of dataset-provided speaker IDs, not the ID counts released with each corpus; crowdsourced sources (Common Voice, Bengali.AI Speech) assign IDs per registration, so a single voice may span many IDs. Totals are corpus-level figures from § 2 .
Abbrev.
Full expansion (Bangla)
\banglafont সাঃ
\banglafont সাল্লাল্লাহু আলাইহি ওয়া সাল্লাম
\banglafont রাঃ
\banglafont রাদিআল্লাহু আনহু
\banglafont রহঃ
\banglafont রাহিমাহুল্লাহ
\banglafont হাঃ
\banglafont হাফিজাহুল্লাহ
\banglafont মাঃ
\banglafont মাদ্দাজিল্লুহু
\banglafont দাঃ বাঃ
\banglafont দামাত বারাকাতুহুম
Appendix
Table 24: Representative entries from the 47-term Bangla honorific dictionary (Stage 2, step 5).
Term
Definition
ARR
ACL Rolling Review submission cycle.
AR
Autoregressive (transformer decoder).
BPE
Byte-Pair Encoding (subword tokenization).
CER
Character Error Rate (ASR transcription error).
CI
95% Confidence Interval.
ECAPA-TDNN
Emphasized Channel Attention, Propagation and Aggregation TDNN.
Appendix
Table 25: Abbreviations, symbols, and Bangla terms used in the paper.
Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving long-form prosody and consistent single-speaker narration uncovered. We present BanglaKontho, a single-speaker Bangla TTS corpus of 20 hours derived from professional audiobook recordings: 7,050 segmented utterances with verified transcripts at 24 kHz. We also release a reusable Bangla text normalizer covering Bangladeshi-style digit grouping, currency and date expressions, Danda punctuation and Unicode normalization, together with the full preprocessing pipeline. An MB-iSTFT-VITS baseline trained from scratch reaches 9.5% WER and 4.46 naturalness MOS, against 16.0% and 3.16 for the same architecture retrained on the 12-hour IndicTTS-Bn corpus. The corpus is released openly under CC BY-NC 4.0.
In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using the International Phonetic Alphabet (IPA) as a unified representation. To eliminate the reliance on prompt text without complex preprocessing like forced alignment, we design a two-stage training paradigm. In Stage 1, we establish X-Voices1 through standard conditional flow-matching training and use it to synthesize 10K hours of speaker-consistent segments as audio prompts. In Stage 2, we fine-tune on these audio pairs with prompt text masked to derive X-Voices2, which enables zero-shot voice cloning without requiring transcripts of audio prompts. Architecturally, we extend F5-TTS by implementing a dual-level injection of language identifiers and decoupling and scheduling of Classifier-Free Guidance to facilitate multilingual speech synthesis. Subjective and objective evaluation results demonstrate that X-Voice outperforms existing flow-matching based multilingual systems like LEMAS-TTS and achieves zero-shot cross-lingual cloning capabilities comparable to billion-scale models such as Qwen3-TTS. To facilitate research transparency and community advancement, we open-source all related resources.
Rixi Xu, Qingyu Liu, Haitao Li +10
1MoE Key Lab of Artificial Intelligence, X-LANCE Lab, Shanghai Jiao Tong University · Center for Language and Speech Processing, Johns Hopkins University · 2Shanghai Innovation Institute +4
We present JaiTTS-v1.0, a state-of-the-art Thai voice cloning text-to-speech model built through continual training on a large Thai-centric speech corpus. The model architecture is adapted from VoxCPM, a tokenizer-free autoregressive TTS model. JaiTTS-v1.0 directly processes numerals and Thai-English code-switching, which is very common in realistic settings, without explicit text normalization. We test the models on short- and long-duration speech generation, which reflects many real-world use cases. JaiTTS-v1.0 achieves a state-of-the-art CER of 1.94%, surpassing the human ground truth of 1.98% for short-duration tasks while performing on par with human ground truth for long-duration tasks. In human judgment evaluations, our model wins 283 of 400 pairwise comparisons against commercial flagships, with only 58 losses. Our code and demo are available at https://github.com/JTS-AI-Team/JaiTTS .