Long-form text-to-speech must retain requested words over extended generations, yet sentence-level training and evaluation can hide omissions and early stops. We introduce Balalaika-Longform, an open Russian corpus of 189 hours in continuous units of 30 seconds to 15 minutes. Long units and matched short windows support fine-tuning comparisons on the same source recordings, and the accompanying evaluation retains every synthesis attempt. We fine-tune CosyVoice3, Qwen3-TTS, VoxCPM2 and F5-TTS on either view and test continuous synthesis on 50 text-voice pairs with voices unseen in fine-tuning, at about 75, 300 and 1,200 words. At 1,200 words, CosyVoice3 WER falls from 99.9% after short-window fine-tuning to 47.3% with long targets, and to 16.6% with punctuated, format-matched transcripts; paired intervals favor long targets for Qwen3-TTS and VoxCPM2, and by a small margin for F5-TTS, whose absolute WER stays above 90%. The results isolate the effect of training sequence length under a fixed continuous-generation protocol.
Figures & tables
Split
Segments
Hours
Videos
Channels
Train
6,043
170.68
265
27
Dev
120
8.50
17
7
Test
193
9.07
30
7
Corpus
6,380
189.13
313
41
Table 1: Channel-disjoint splits before training-text exclusions. The remaining 24 segments (0.88 h) come from seven held-out videos, one otherwise unrepresented, that fail the ASR-consistency gate; they are released separately and never used.
Complete ↑ (%)
All-attempt WER ↓ (%)
Bucket
Target duration
Base
Short-SFT
Long-SFT
Long-SFT-Punct
Base
Short-SFT
Long-SFT
Long-SFT-Punct
B0
20–40 s
83.3
91.7
100.0
100.0
7.18
7.07
3.45
2.93
B1
60–90 s
58.3
41.7
83.3
75.0
14.56
17.04
11.70
9.61
B2
2–3 min
0.0
0.0
100.0
100.0
90.86
82.64
2.74
3.42
B3
4–6 min
0.0
0.0
41.7
75.0
98.07
97.77
47.22
36.66
B4
8–12 min
0.0
0.0
33.3
66.7
99.82
99.89
37.22
19.15
Table 2: CosyVoice3 results on the earlier two-voice set (six texts, five duration buckets, two held-out voices; 60 attempts per arm) by input duration; all synthesis attempts are retained. Best value per row in bold.
Figure 1: Windowed diagnostics on the same 22 longest inputs per condition; Base and Long-SFT-Punct are shown, and all 13 conditions accompany the code release. Both metrics use 5-s windows with a 2.5-s hop; speaker windows require at least 1 s of voiced audio. Faint curves are item-weighted means, bold curves exponential smoothing ( α=0.3 ); lower strips count contributing items (at least two per score).
All-attempt WER ↓ (%) [95% CI] by input length
Recall ↑ (%)
Backbone
Arm
∼ 75 words
∼ 300 words
∼ 1,200 words
∼ 1,200 words
CosyVoice3
Base
7.3 [4.1, 11.2]
84.8 [77.9, 90.9]
99.8 [99.7, 99.9]
0.2
Short-SFT
7.7 [4.9, 10.8]
82.2 [72.5, 90.3]
99.9 [99.9, 100.0]
0.1
Long-SFT
5.2 [2.8, 8.0]
13.8 [5.7, 24.1]
47.3 [37.2, 57.7]
69.5
Long-SFT-Punct
3.6 [2.2, 5.0]
5.6 [2.7, 9.8]
16.6 [10.5, 23.2]
85.9
Qwen3-TTS
Base
3.9 [2.0, 6.1]
5.3 [2.5, 8.6]
66.9 [62.8, 70.8]
34.4
Table 3: Published-text evaluation of all 13 continuous conditions by input length (14, 14 and 22 attempts per condition; texts from 22 works, 21 voices unseen in fine-tuning). All-attempt WER and correct-word recall include failed outputs; brackets are work-cluster bootstrap 95% intervals; text differences use unrounded means; best point estimate per backbone and column in bold (not a significance statement). Qwen3-TTS, VoxCPM2 and F5-TTS also use punctuated text for Short-SFT.