Long-form text-to-speech must retain requested words over extended generations, yet sentence-level training and evaluation can hide omissions and early stops. We introduce Balalaika-Longform, an open Russian corpus of 189 hours in continuous units of 30 seconds to 15 minutes. Long units and matched short windows support fine-tuning comparisons on the same source recordings, and the accompanying evaluation retains every synthesis attempt. We fine-tune CosyVoice3, Qwen3-TTS, VoxCPM2 and F5-TTS on either view and test continuous synthesis on 50 text-voice pairs with voices unseen in fine-tuning, at about 75, 300 and 1,200 words. At 1,200 words, CosyVoice3 WER falls from 99.9% after short-window fine-tuning to 47.3% with long targets, and to 16.6% with punctuated, format-matched transcripts; paired intervals favor long targets for Qwen3-TTS and VoxCPM2, and by a small margin for F5-TTS, whose absolute WER stays above 90%. The results isolate the effect of training sequence length under a fixed continuous-generation protocol.
Figures & tables
Split
Segments
Hours
Videos
Channels
Train
6,043
170.68
265
27
Dev
120
8.50
17
7
Test
193
9.07
30
7
Corpus
6,380
189.13
313
41
Table 1: Channel-disjoint splits before training-text exclusions. The remaining 24 segments (0.88 h) come from seven held-out videos, one otherwise unrepresented, that fail the ASR-consistency gate; they are released separately and never used.
Complete ↑ (%)
All-attempt WER ↓ (%)
Bucket
Target duration
Base
Short-SFT
Long-SFT
Long-SFT-Punct
Base
Short-SFT
Long-SFT
Long-SFT-Punct
B0
20–40 s
83.3
91.7
100.0
100.0
7.18
7.07
3.45
2.93
B1
60–90 s
58.3
41.7
83.3
75.0
14.56
17.04
11.70
9.61
B2
2–3 min
0.0
0.0
100.0
100.0
90.86
82.64
2.74
3.42
B3
4–6 min
0.0
0.0
41.7
75.0
98.07
97.77
47.22
36.66
B4
8–12 min
0.0
0.0
33.3
66.7
99.82
99.89
37.22
19.15
Table 2: CosyVoice3 results on the earlier two-voice set (six texts, five duration buckets, two held-out voices; 60 attempts per arm) by input duration; all synthesis attempts are retained. Best value per row in bold.
Figure 1: Windowed diagnostics on the same 22 longest inputs per condition; Base and Long-SFT-Punct are shown, and all 13 conditions accompany the code release. Both metrics use 5-s windows with a 2.5-s hop; speaker windows require at least 1 s of voiced audio. Faint curves are item-weighted means, bold curves exponential smoothing ( α=0.3 ); lower strips count contributing items (at least two per score).
All-attempt WER ↓ (%) [95% CI] by input length
Recall ↑ (%)
Backbone
Arm
∼ 75 words
∼ 300 words
∼ 1,200 words
∼ 1,200 words
CosyVoice3
Base
7.3 [4.1, 11.2]
84.8 [77.9, 90.9]
99.8 [99.7, 99.9]
0.2
Short-SFT
7.7 [4.9, 10.8]
82.2 [72.5, 90.3]
99.9 [99.9, 100.0]
0.1
Long-SFT
5.2 [2.8, 8.0]
13.8 [5.7, 24.1]
47.3 [37.2, 57.7]
69.5
Long-SFT-Punct
3.6 [2.2, 5.0]
5.6 [2.7, 9.8]
16.6 [10.5, 23.2]
85.9
Qwen3-TTS
Base
3.9 [2.0, 6.1]
5.3 [2.5, 8.6]
66.9 [62.8, 70.8]
34.4
Table 3: Published-text evaluation of all 13 continuous conditions by input length (14, 14 and 22 attempts per condition; texts from 22 works, 21 voices unseen in fine-tuning). All-attempt WER and correct-word recall include failed outputs; brackets are work-cluster bootstrap 95% intervals; text differences use unrounded means; best point estimate per backbone and column in bold (not a significance statement). Qwen3-TTS, VoxCPM2 and F5-TTS also use punctuated text for Short-SFT.
We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER consensus decoding, while retaining optional word-level timestamps, followed by automatic quality and speaker-purity filtering. The text is further enriched with punctuation restoration, lexical stress and "\textipa{e}/\textipa{He}" normalization, and IPA phonemes. Using Balalaika, we build a 5.1k-hour multi-source Russian corpus with rich annotations, and show consistent gains under equalized training budgets for both speech denoising and TTS; ablations confirm complementary benefits of stress and punctuation and improved synthesis with stricter MOS filtering. The datasets are publicly available at \href{https://huggingface.co/collections/lab260/balalaika-dataset}{\underline{\textbf{HuggingFace}}}
Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence boundary artifacts. Existing approaches either compress sequences, increase context length or naively concatenate independently synthesized chunks. We present an inference-time approach called MagpieTTS-LF that enables MagpieTTS to produce coherent long-form speech without model retraining. Our method introduces three key innovations: (1) soft attention priors to guide monotonic alignment while preserving past and future context; (2) a stateful inference algorithm that maintains context across sentence chunks, ensuring prosodic continuity; (3) history-aware text encoding that uses past text for discourse-level prosodic planning. Experiments on long texts show significant improvements in long-range intelligibility, prosodic coherence, speaker consistency, and boundary naturalness compared to other baselines.
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and speaker similarity (SIM) on short-form prompts but significantly deteriorate when used with long-form prompts. We propose Localized Attention-Constrained Inference (LACI), an inference-only method to detect TTS errors in near real-time, roll back to the error onset and regenerate with temporary guardrails, adding negligible computational overhead. Using LACI, we improve worst-of-N WER across 10 RNG seeds for Qwen3-TTS-0.6B from 35.2% to 3.4% on prompts longer than 1500 words, even surpassing its short-form reliability of 5.4% on prompts with fewer than 500 words. To demonstrate the efficacy of LACI on voice cloning reliability, we propose a sliding-window version of the SIM metric that we call wSIM. wSIM exposes several novel failure patterns that are not captured by SIM. LACI improves worst-of-N wSIM from 0.01 to 0.47 on 120 seconds of reference audio while reducing the rate of catastrophic generations with WER above 30% from 26% to below 1%
Rongxiang Wang, Berkin Durmus, Aysegul Orhon +2
Argmax, Inc. · University of Virginia · Bilkent University