Instruction text-to-speech (ITTS) systems systematically encode acoustic gender skews from descriptive style prompts, such as occupations or personas, even when demographic attributes are left unspecified. Calibrating the implicit gender distribution of the synthesized voices, across heterogeneous architectures and without model retraining or overriding explicit user prompts, is an open problem. In this work, we propose \textit{model-adaptive steering}, a training-free bias calibration method that steers post-encoder conditioning representations using a group deviation vector paired with a coarse-to-fine strength search on a development set. A deterministic lexical gate bypasses intervention whenever explicit gender keywords are detected, preserving intended prompt semantics. Evaluated on a 12,800-prompt held-out benchmark across four ITTS models (three architectures), the proposed method reduces aggregate calibration error from 11.5--28.1 to 0.8--5.9 percentage points (e.g., shifting Parler-TTS Mini from 78.1% and PromptTTS++ from 27.3% female to 50.8--52.4%). Output rates stay within 0.9 pp across anchor set sizes at fixed operating points, while UTMOS decreases by at most 0.08 and WER degrades by at most 3.3 pp.
Figures & tables
Figure 1: Two-stage calibration and inference framework. Stage 1 estimates the group deviation vector d from anchor pairs AN and identifies the operating point θ∗ on Ddev via coarse-to-fine search (Sec. 2.2 ). Stage 2 shifts post-encoder conditioning representations H prior to acoustic decoding, where a deterministic lexical gate bypasses steering whenever explicit gender keywords are detected.
Female Classification Rate (%)
Calibration Error (pp) ↓
Utility
Model
Setting
All
Status
Career
Persona
2-axis
3-axis
Eagg
Emacro
Eworst
UTMOS ↑
WER ↓
Parler-Mini
Original
78.11
82.00
79.42
76.92
78.77
77.71
28.11
28.97
32.00
3.8346
0.1826
Neutral
86.07
76.00
85.08
88.13
85.26
85.45
36.07
33.98
38.13
3.9021
0.1314
LEACE pb ( β=1.00 )
74.84
77.00
83.08
67.33
78.03
74.13
24.84
25.91
33.08
3.8512
0.1419
Ours ( α=2.5 )
52.43
51.00
56.62
46.79
54.94
53.55
2.43
3.86
6.62
3.8059
0.1413
Parler-Large
Original
74.83
53.00
88.46
85.69
63.99
61.29
24.83
20.49
38.46
3.5362
0.2487
Table 1: Held-out evaluation ( 12,800 ) across prompt strata and calibration errors. Error metrics are reported in percentage points (pp), with pF in % ( Eagg weights strata by size; Emacro weights them equally): Eagg=∣pF−50∣ , Emacro=51∑s∣pF,s−50∣ , and Eworst=maxs∣pF,s−50∣ ( ↓ lower is better). Speech utility (UTMOS ↑ , WER ↓ ) is evaluated on a shared 500-prompt test subset (WER as a fraction). 95% bootstrap CI half-widths: ≤ 0.88 pp (All), ≤ 2.00 pp ( Emacro ); 7 utterances < 0.1 s are excluded. Test strata hold 100/2,600/3,900/3,100/3,100 prompts, thus status-driven Eworst carries ± 9.8 pp. All representation-based methods are fitted on the same 2,500 anchor pairs. Bold denotes the best debiasing result within each model; underline denotes the second best (Original serves as the unmitigated baseline).
Model
−2.5
−2.0
−1.5
−1.0
0.0
1.0
1.5
2.0
2.5
Parler-Mini
96.8
96.6
94.4
90.6
78.6
66.2
62.4
55.0
53.2 ∗
Parler-Large
89.0
86.8
84.4
83.4
73.4
65.8
59.2
51.0 ∗
44.4
PromptTTS++
99.6
100.0
100.0
100.0
26.2
0.0
0.0
0.0
0.0
Table 2: Female classification rates (%) on the 500-item development set. Bold with an asterisk (*) indicates the selected operating point θ∗ .
Anchor Pairs ( N )
Model
500
1,500
2,500
Δ (pp)
Parler-Mini
52.60
52.82
52.43
0.39
Parler-Large
52.00
52.19
51.31
0.88
PromptTTS++
51.20
50.50
50.80
0.70
VoxInstruct
56.30
56.06
55.86
0.44
Table 3: Sensitivity of female classification rates (%) to anchor pair count N on the held-out test set ( 12,800 ). Calibrated operating points θ∗ remain fixed. Δ denotes the maximum percentage point variation across pair counts.