Instruction text-to-speech (ITTS) systems systematically encode acoustic gender skews from descriptive style prompts, such as occupations or personas, even when demographic attributes are left unspecified. Calibrating the implicit gender distribution of the synthesized voices, across heterogeneous architectures and without model retraining or overriding explicit user prompts, is an open problem. In this work, we propose \textit{model-adaptive steering}, a training-free bias calibration method that steers post-encoder conditioning representations using a group deviation vector paired with a coarse-to-fine strength search on a development set. A deterministic lexical gate bypasses intervention whenever explicit gender keywords are detected, preserving intended prompt semantics. Evaluated on a 12,800-prompt held-out benchmark across four ITTS models (three architectures), the proposed method reduces aggregate calibration error from 11.5--28.1 to 0.8--5.9 percentage points (e.g., shifting Parler-TTS Mini from 78.1% and PromptTTS++ from 27.3% female to 50.8--52.4%). Output rates stay within 0.9 pp across anchor set sizes at fixed operating points, while UTMOS decreases by at most 0.08 and WER degrades by at most 3.3 pp.
Figures & tables
Figure 1: Two-stage calibration and inference framework. Stage 1 estimates the group deviation vector d from anchor pairs AN and identifies the operating point θ∗ on Ddev via coarse-to-fine search (Sec. 2.2 ). Stage 2 shifts post-encoder conditioning representations H prior to acoustic decoding, where a deterministic lexical gate bypasses steering whenever explicit gender keywords are detected.
Female Classification Rate (%)
Calibration Error (pp) ↓
Utility
Model
Setting
All
Status
Career
Persona
2-axis
3-axis
Eagg
Emacro
Eworst
UTMOS ↑
WER ↓
Parler-Mini
Original
78.11
82.00
79.42
76.92
78.77
77.71
28.11
28.97
32.00
3.8346
0.1826
Neutral
86.07
76.00
85.08
88.13
85.26
85.45
36.07
33.98
38.13
3.9021
0.1314
LEACE pb ( β=1.00 )
74.84
77.00
83.08
67.33
78.03
74.13
24.84
25.91
33.08
3.8512
0.1419
Ours ( α=2.5 )
52.43
51.00
56.62
46.79
54.94
53.55
2.43
3.86
6.62
3.8059
0.1413
Parler-Large
Original
74.83
53.00
88.46
85.69
63.99
61.29
24.83
20.49
38.46
3.5362
0.2487
Table 1: Held-out evaluation ( 12,800 ) across prompt strata and calibration errors. Error metrics are reported in percentage points (pp), with pF in % ( Eagg weights strata by size; Emacro weights them equally): Eagg=∣pF−50∣ , Emacro=51∑s∣pF,s−50∣ , and Eworst=maxs∣pF,s−50∣ ( ↓ lower is better). Speech utility (UTMOS ↑ , WER ↓ ) is evaluated on a shared 500-prompt test subset (WER as a fraction). 95% bootstrap CI half-widths: ≤ 0.88 pp (All), ≤ 2.00 pp ( Emacro ); 7 utterances < 0.1 s are excluded. Test strata hold 100/2,600/3,900/3,100/3,100 prompts, thus status-driven Eworst carries ± 9.8 pp. All representation-based methods are fitted on the same 2,500 anchor pairs. Bold denotes the best debiasing result within each model; underline denotes the second best (Original serves as the unmitigated baseline).
Model
−2.5
−2.0
−1.5
−1.0
0.0
1.0
1.5
2.0
2.5
Parler-Mini
96.8
96.6
94.4
90.6
78.6
66.2
62.4
55.0
53.2 ∗
Parler-Large
89.0
86.8
84.4
83.4
73.4
65.8
59.2
51.0 ∗
44.4
PromptTTS++
99.6
100.0
100.0
100.0
26.2
0.0
0.0
0.0
0.0
Table 2: Female classification rates (%) on the 500-item development set. Bold with an asterisk (*) indicates the selected operating point θ∗ .
Anchor Pairs ( N )
Model
500
1,500
2,500
Δ (pp)
Parler-Mini
52.60
52.82
52.43
0.39
Parler-Large
52.00
52.19
51.31
0.88
PromptTTS++
51.20
50.50
50.80
0.70
VoxInstruct
56.30
56.06
55.86
0.44
Table 3: Sensitivity of female classification rates (%) to anchor pair count N on the held-out test set ( 12,800 ). Calibrated operating points θ∗ remain fixed. Δ denotes the maximum percentage point variation across pair counts.
Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.g., crosstalk, dialects). Moreover, fine-tuning pretrained models requires large, high-quality datasets, limiting rapid personalization. We propose VoiceTTA, a reinforcement learning-based test-time adaptation (TTA) method that improves voice imitation of pretrained zero-shot TTS models. VoiceTTA introduces two style rewards based on coefficient-of-variation differences of F0 and energy, combined with speaker similarity and intelligibility (WER from a pretrained Whisper model), and optimizes learnable prefixes via group relative preference optimization (GRPO) in a flow matching-based model at inference time. Extensive experiments demonstrate substantial improvements on uncommon speech prompts, outperforming state-of-the-art baselines. Audio samples are available at https://voicetta.pages.dev/
Tianxin Xie, Chenxing Li, Dong Yu +1
The Hong Kong University of Science and Technology (Guangzhou) · Tencent
Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithful system should treat a speaker as who they sound like, not as whoever usually says what they said. Testing this is harder than it looks, since most S2S models answer in a single, fixed output voice, hard-coded so it cannot drift toward a stereotype. Checking the output voice comes back clean even when the model is biased. We therefore ask two questions. When a model re-speaks the input, does the stereotype in the words shift the perceived gender of the output voice (voice rendering)? And when the model states the speaker's gender, does it follow the voice or the content (gender attribution)? We answer both with one controlled experiment crossing male and female voices with masculine-, neutral-, and feminine-stereotyped passages, on five open- and closed-source models in English, Spanish, and Mandarin. The rendered voice shows no stereotype drift. But every model decides the speaker's gender from the content, not the voice. Making the content one step more feminine (masculine -> neutral -> feminine) multiplies the odds of a "female" judgment by 1.7-24. When the content clashes with the voice, the worst model misgenders the speaker in 90% of cases. When they agree, it misgenders in only 2%. The bias thus hides in gender attribution, where fixed-voice evaluation cannot see, and where audits must look as S2S systems increasingly speak for real people.
Instruct-TTS systems expand structured style labels into natural-language training instructions through LLM rewriting, yet we find that over 40% of unconstrained rewrites contain semantic drift that corrupts supervision and weakens generalization. We formalize this problem as instruction supervision instability and propose a data-centric stabilization recipe that jointly improves coverage and fidelity through three mechanisms: controllable instruction diversification for systematic expansion, LLM-based drift filtering for quality control, and attribute-aligned supervision that grounds prosody control in acoustic perturbations. On the Chinese split of InstructTTSEval, our recipe raises instruction-following from 34.5% without fine-tuning and 51.0% with naive fine-tuning to 56.4%, while constrained rewriting reduces drift from 40.4% to 15.4%. Ablations confirm the three mechanisms are complementary, and the drift taxonomy may generalize to instruction-driven generation beyond TTS.
Yizhong Geng, Kecan Mao, Qifei Li +6
Beijing University of Posts and Telecommunications, China · Li Auto, China