Instruction-based text-to-speech (TTS) offers control over voice characteristics and speech expression through interfaces including voice cloning and text-based voice design. Voice cloning reproduces a reference voice, whereas text-based voice design creates a voice from a natural-language description. However, neither interface directly enables users to modify the timbre of a given reference and synthesize speech with the modified voice. Meanwhile, utterance-level expressive instructions leave changes across individual text segments underspecified. We introduce \textbf{EDICT}, a framework that unifies global timbre editing and local expressive control by using an edited acoustic reference to anchor voice identity across segments. To enable synthesis with an instruction-edited voice, EDICT combines reference audio with structured timbre edits to generate an edited reference in codec-token space. This representation serves as a shared voice anchor for a frozen TTS backbone, allowing segment-specific natural-language instructions to guide expression. To accommodate instruction changes while supporting acoustic continuity, EDICT rebuilds the KV cache at each segment boundary, refreshing instruction conditioning while retaining bounded acoustic context from previously generated speech. Evaluations on our proposed TimbreEdit-Bench and IntraTTS-Bench demonstrate improved timbre editing and a favorable balance between local instruction adherence, speaker consistency, and transition quality. Audio demos are available.
Figures & tables
Figure 1: EDICT combines global voice editing with local delivery control. The trainable editor uses the source voice and global edit instruction to generate an edited codec reference. This reference conditions every segment of the frozen TTS backbone, whose KV cache is rebuilt at instruction boundaries using selected acoustic context. The segments form one continuous utterance.
Figure 2: Segment-boundary KV cache reconstruction. Each non-final <EOS> triggers a fresh prefill containing the next instruction and text, the shared edited reference, and acoustic inputs from the first S and most recent K generated positions. The upper-left panel shows replacement of the old cache; the upper-right panel shows causal attention after reconstruction. The lower timeline illustrates successive instruction switches within one codec stream.
Figure 3: Benchmark construction and evaluation. Left: TimbreEdit-Bench pairs source and target voices with editing instructions. Right: IntraTTS-Bench aligns delivery instructions with successive text segments.
System
Instruction
Content and quality
WavLM
ERes2Net
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-T ↑
SIM-R
Δ↑
SIM-T ↑
SIM-R
Δ↑
Instruction-conditioned systems
Qwen3-TTS-VD
Structured
4.5 / 0.3
3.861
3.946
0.701
0.871
−0.170
0.450
0.745
−0.295
Descriptive
3.0 / 0.3
3.830
3.965
0.676
0.832
−0.156
0.435
0.714
−0.279
Qwen3-TTS-Base
Structured
5.1 / 1.9
3.742
3.805
0.632
0.950
−0.318
0.431
0.845
−0.414
Descriptive
3.3 / 1.1
3.756
3.776
0.637
0.950
−0.313
0.434
0.849
−0.415
Table 1: Objective results on TimbreEdit-Bench. WER/CER reports English WER / Chinese CER in percent. SIM-T/R denotes target/source similarity; Δ=SIM-T−SIM-R . Best and second-best scores within each instruction form are bold and underlined , with CER and WER ranked separately; SIM-R is unranked. Target-reference voice conversion receives target audio and is evaluated separately on a same-text subset, with its best similarity and Δ scores in bold .
System
Instruction
Up Acc. ↑
Down Acc. ↑
Category Acc. ↑
Preservation ↑
Qwen3-TTS-VD
Structured
67.0
64.3
53.2
83.3
Descriptive
60.0
63.0
49.8
85.0
Qwen3-TTS-Base
Structured
62.3
54.5
46.4
45.7
Descriptive
66.0
54.5
50.0
45.6
CosyVoice 2
Structured
68.1
65.5
57.7
68.2
Descriptive
72.2
73.3
69.0
66.2
Table 2: Attribute-operation accuracy (%) on 500 TimbreEdit-Bench requests. Gemini compares source and generated audio without the target recording. Up/Down assesses ordinal direction, Category assesses categorical targets, and Preservation assesses unchanged attributes. Best and second-best scores within each instruction form are bold and underlined .
System
Control (%)
Consistency
Content and quality
SegAcc ↑
EmoAcc ↑
Trans ↑
SIM-I ↑
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
Qwen3-TTS-VD (Concat.)
92.44
92.66
60.68
0.658
4.5 / 0.9
3.266
3.939
Qwen3-TTS-VD (Joint)
66.80
66.37
70.66
0.922
3.8 / 1.3
3.794
3.952
Qwen3-TTS-VD (TED-TTS)
74.25
74.02
78.03
0.911
4.5 / 1.7
3.655
3.853
Qwen3-TTS-VD (EDICT-Stage II)
82.30
83.60
82.00
0.915
5.1 / 1.8
3.631
3.930
CosyVoice 2 (Concat.)
62.43
62.73
31.80
0.917
6.5 / 3.6
3.916
3.866
Table 3: Intra-utterance instruction control on IntraTTS-Bench. SegAcc, EmoAcc, and Trans measure instruction adherence, emotion realization, and transition quality in percent. SIM-I measures inter-segment speaker similarity; WER/CER reports English WER / Chinese CER in percent. Best and second-best scores within each backbone are bold and underlined , with CER and WER ranked separately.
Configuration
Control (%)
Content and quality
Consistency
SIM-T ↑
SegAcc ↑
Trans ↑
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-I ↑
WavLM
ERes2Net
Qwen3-TTS-VD (Joint)
60.19
70.07
3.6 / 1.2
3.339
3.869
0.918
0.683
0.395
Qwen3-TTS-VD (Concat.)
86.43
56.69
4.5 / 0.8
3.669
3.895
0.884
0.694
0.405
EDICT Stage I + Concat.
91.17
58.47
4.8 / 1.1
3.296
3.739
0.881
0.749
0.598
EDICT
81.62
83.45
5.1 / 1.9
3.795
3.921
0.917
0.788
0.609
Table 4: Joint timbre and delivery control with Qwen3-TTS-VD. Requests combine TimbreEdit-Bench voice edits with IntraTTS-Bench texts and delivery instructions. SIM-T compares the complete output with the target voice; SIM-I measures inter-segment speaker similarity. WER/CER reports English WER / Chinese CER in percent. Best and second-best scores are bold and underlined , with CER and WER ranked separately.
Figure 4: Subjective evaluation of timbre editing and local control. Bars show mean five-point ratings from 19 listeners; error bars are computed from 95% bootstrap confidence intervals over audio items. Each system has 10 items in (a) and 12 in (b). EDICT is highlighted in coral; all variants in (b) use Qwen3-TTS-VD. † Seed-VC receives target audio and is not rated for instruction-based edit accuracy.
Configuration
Instruction
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-T ↑
WavLM
ERes2Net
SFT (0.5M)
Structured
3.9 / 1.5
3.689
3.895
0.736
0.498
Descriptive
4.2 / 1.5
3.708
3.885
0.698
0.481
SFT (1M)
Structured
4.1 / 0.9
3.936
3.896
0.757
0.561
Descriptive
3.9 / 0.9
3.666
3.913
0.693
0.556
DPO ( M=4 )
Structured
5.1 / 1.5
3.885
3.962
0.761
0.552
Table 5: Timbre-editor training ablations on TimbreEdit-Bench. All post-training variants start from SFT (1M). M denotes the candidate count per condition; wDPO includes auxiliary SFT as defined in Equation 2 . SIM-T measures target-voice similarity; WER/CER reports English WER / Chinese CER in percent. Best and second-best scores within each instruction form are bold and underlined , with CER and WER ranked separately.
Configuration
Control (%)
Consistency
Content and quality
SegAcc ↑
EmoAcc ↑
Trans ↑
SIM-I ↑
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
Persistent KV cache
54.25
54.02
68.00
0.918
3.9 / 0.8
3.555
3.953
Truncated KV cache
58.05
56.70
64.94
0.892
4.4 / 1.1
3.278
3.905
EDICT-Stage II
82.30
83.60
82.00
0.915
5.1 / 1.8
3.631
3.930
w/o initial context
89.29
89.00
79.99
0.863
4.3 / 1.4
3.585
3.896
w/o recent context
81.15
82.26
55.32
0.875
4.2 / 1.1
3.548
3.830
Table 6: KV cache reconstruction ablations on IntraTTS-Bench. All variants use Qwen3-TTS-VD. The full configuration retains the first S=4 and most recent K=5 acoustic positions; the context ablations remove one of these groups. Persistent and truncated caches reuse old KV states, whereas EDICT recomputes them. Best and second-best scores are bold and underlined , with WER and CER ranked separately.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Control input
Control target
Scope
Speech content editing
Recording + revised text
Spoken content
Selected span
Timbre editing
Reference + edit instruction
Selected timbre attributes
Utterance
Voice cloning
Reference speech
Voice identity
Utterance
Voice design
Description / attribute specification
Voice characteristics
Utterance
Global expressive control
Utterance-level instruction
Emotion, rate, prosody, etc
Utterance
Local expressive control
Span-specific instructions
Emotion, rate, prosody, etc
Word / segment
Appendix
Table 7: Taxonomy of speech control tasks. Typical control inputs, control targets, and temporal scope are shown. Inputs are not exhaustive, and tasks may overlap.
Method
Zero-shot cloning
Voice design
Timbre modification
Intra-utterance control
Control formulation
VoxInstruct ( Zhou et al., 2024 )
✓
✓
✗
✓ a
Instruction-guided synthesis and word-level emphasis
Qwen3-TTS-Base ( Hu et al., 2026a )
✓
✗
✗
✗
Reference-based voice cloning
Qwen3-TTS-VoiceDesign ( Hu et al., 2026a )
✓
✓
✗
✗
Description-guided voice design
Spark-TTS ( Wang et al., 2025b )
✓
✓
✗
✗
Attribute-guided design and reference cloning
VoiceSculptor ( Hu et al., 2026b )
✓
✓
✗
✗
Voice design followed by cloning
CosyVoice 2 ( Du et al., 2024 )
✓
✗
✗
✓ b
Instructed cloning and phrase-level tags
Appendix
Table 8: Comparison of speech generation and editing methods. Capabilities refer to the named models in the cited settings. Superscripts specify restricted scope or additional training.
Attribute
Values
Perceptual description
Perceived gender
Masculine, Feminine
Perceived vocal masculinity or femininity.
Age impression
Child, Young adult, Mature
Perceived age conveyed by the voice.
Pitch register
Low, Mid, High
Typical pitch range of the voice.
Brightness
Dark, Neutral, Bright
Perceived prominence of high-frequency energy.
Vocal weight
Thin, Medium, Thick
Perceived fullness and body of the voice.
Resonance focus
Head, Balanced, Chest
Resonance perceived as head- or chest-focused.
Appendix
Table 9: Perceptual attribute schema for timbre editing. Perceived gender and age impression are categorical; the seven remaining attributes use ordinal levels 1–3 in the listed order.
Figure 5: Synthetic speech-bank statistics. (a) Retained recording durations in 0.5-second bins. (b,c) Perceived gender and age impression. (d) Distributions of the seven ordinal attributes; levels 1–3 follow the schema in Table 9 . Profile percentages are normalized within the training (5,800) and held-out (200) subsets. Young denotes young adult.
Figure 6: Edit coverage across one million training pairs. (a) Number of changed attributes per pair. (b,c) Source-to-target transitions for perceived gender and age impression. (d) Signed target-minus-source differences for ordinal attributes; Keep denotes no change. Heatmaps use a shared color scale and show percentages of all pairs. Each complete matrix in (b,c) and each column in (d) sums to 100% before rounding. Young denotes young adult.
Figure 7: IntraTTS-Bench composition. (a,b) Request counts by segment count and topic, split by language. (c,d) Segment lengths in bins of width two, normalized within each language. Chinese lengths count non-whitespace characters, including punctuation; English lengths count whitespace-separated words.
Criterion
Valid annotations
Agreement (%)
H1
H2
Gemini
H1–H2
H1–Gemini
H2–Gemini
SegAcc
197
198
200
90.4
83.8
85.9
EmoAcc
197
198
200
85.8
83.2
85.4
Trans
99
100
100
91.9
93.9
96.0
Appendix
Table 10: Human–Gemini agreement on intra-utterance control. Valid counts refer to segment-level judgments for SegAcc and EmoAcc, and boundary-level judgments for Trans. Agreement (%) uses jointly valid items for each pair.
Criterion
Valid annotations
Agreement (%)
H1
H2
Gemini
H1–H2
H1–Gemini
H2–Gemini
Up
100
100
100
81.0
79.0
76.0
Down
100
100
100
73.0
67.0
66.0
Category
100
100
100
82.0
84.0
78.0
Preservation
100
100
100
74.0
72.0
70.0
Appendix
Table 11: Human–Gemini agreement on timbre editing. Each annotator provides 100 valid judgments per criterion. Agreement (%) is computed from paired pass/fail labels.
Figure 8: Instructions shown to listening-test participants. Panel (a) gives general guidance on playback and rating; panels (b) and (c) describe the timbre-editing and intra-utterance control tasks, respectively. System identities are hidden.
Figure 9: Listening-test interfaces. Panel (a) collects naturalness, edit accuracy, and target timbre similarity ratings in sequence; the target reference is revealed only after edit accuracy is rated. Panel (b) collects naturalness, instruction following, and transition quality ratings for the complete utterance. All criteria use five-point scales.
Figure 10: Attention across segment transitions. Rows show zero-indexed Transformer layers 7 and 14; columns compare TED-TTS and EDICT on the same request, using separate rollouts. We average attention over all 16 heads in each layer. G denotes the global instruction, Ik/xk the segment instruction/text, H the preceding audio, S/R the initial/recent context, and c2 the new segment’s audio positions. Gray cells are masked or unavailable. All panels share a power color scale with γ=0.35 .
Figure 11: Attention during KV cache reconstruction. Each layer group shows replay with (a) old KV states and old conditions, (b) reconstructed KV states and new conditions, and (c) the full prefill followed by decoding. The dashed coral box marks retained acoustic inputs whose KV states are recomputed; solid lines separate prefill from decoding. Layers 7 and 14 are zero-indexed, with attention averaged over all 16 heads. The replay comparison changes both cache and conditions, so it does not isolate reconstruction alone.
Figure 12: Statistics of the ESD evaluation set. (a) Speaker-pair coverage. (b) Number of edited attributes per request. (c) Attribute-change distributions. In-bar numbers indicate case counts.
System
Instruction
Content and quality
WavLM
ERes2Net
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-T ↑
SIM-R
Δ↑
SIM-T ↑
SIM-R
Δ↑
(a) English
Instruction-conditioned systems
Qwen3-TTS-VD
Structured
5.3
4.219
3.487
0.510
0.714
−0.204
0.383
0.531
−0.148
Descriptive
6.1
4.162
3.400
0.483
0.709
−0.226
0.385
0.533
−0.148
Qwen3-TTS-Base
Structured
4.7
4.358
3.387
0.396
0.934
−0.538
0.301
0.770
−0.469
Appendix
Table 12: Objective results on the ESD evaluation set. WER/CER reports English WER or Mandarin CER. SIM-T/R denotes target/source similarity; Δ=SIM-T−SIM-R . Best and second-best scores within each language and instruction form are bold and underlined ; SIM-R is unranked. Target-reference VC is shown separately, with its best similarity and Δ scores in bold .
Configuration
Instruction
Content and quality
WavLM
ERes2Net
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-T ↑
SIM-T ↑
(a) SFT data scale
SFT (0.5M)
Structured
3.9 / 1.5
3.689
3.895
0.736
0.498
Descriptive
4.2 / 1.5
3.708
3.885
0.698
0.481
SFT (1M)
Structured
4.1 / 0.9
3.936
3.896
0.757
0.561
Descriptive
3.9 / 0.9
3.666
3.913
0.693
0.556
Appendix
Table 13: Extended timbre-editor training ablations on TimbreEdit-Bench. SIM-T measures target-voice similarity; WER/CER reports English WER / Chinese CER in percent. Best and second-best scores within each instruction form are bold and underlined , with CER and WER ranked separately.
Figure 13: Sensitivity to the auxiliary SFT coefficient in wDPO on TimbreEdit-Bench. Solid and dashed curves denote structured and descriptive instructions, respectively. Shading marks the default λSFT=0.01 . The horizontal axis is linear up to 0.01 and logarithmic thereafter.
Initial context S
0
1
2
3
4
5
SIM-I ↑
0.872
0.885
0.902
0.899
0.912
0.915
EmoSim ↑
0.856
0.850
0.842
0.837
0.836
0.821
Appendix
Table 14: Effect of initial acoustic context on a 500-request IntraTTS-Bench subset. With Qwen3-TTS-VD and K=5 , we vary the number S of initial acoustic positions retained during reconstruction. SIM-I measures speaker consistency; EmoSim measures emotion-embedding similarity to independently synthesized reference segments. Bold marks the default configuration, rather than the highest score.
Recent context K
1
2
3
4
5
6
7
8
9
10
SIM-I ↑
0.876
0.895
0.903
0.911
0.912
0.912
0.915
0.917
0.916
0.912
EmoSim ↑
0.837
0.842
0.838
0.832
0.836
0.831
0.825
0.819
0.821
0.817
Appendix
Table 15: Effect of recent acoustic context on a 500-request IntraTTS-Bench subset. With Qwen3-TTS-VD and S=4 , we vary the number K of recent acoustic positions retained during reconstruction. SIM-I measures speaker consistency; EmoSim measures emotion-embedding similarity to independently synthesized reference segments. Bold marks the default configuration, rather than the highest score.
Language
System
WER/CER ↓
Emo2v ↑
SIM-I ↑
TCD ↓
OVRL ↑
UTMOS ↑
English
Concat.
1.18
88.33
0.667
0.467
3.799
3.550
Joint
0.30
59.47
0.929
0.487
3.754
4.040
TED-TTS
0.69
80.91
0.919
0.463
3.676
3.875
EDICT-Stage II
1.09
85.81
0.932
0.357
3.739
3.863
Chinese
Concat.
1.74
85.56
0.589
0.488
3.665
3.528
Joint
1.57
61.48
0.856
0.425
3.692
3.763
Appendix
Table 16: Intra-utterance emotion control on 500 MED-TTS examples. WER/CER denotes English WER or Chinese CER in percent. Emo2v is a segment-level emotion classification accuracy in percent; SIM-I measures inter-segment speaker similarity. TCD is reported as 10−3TCD , and OVRL denotes DNSMOS OVRL. Best and second-best scores within each language are bold and underlined .
Text category
System
WER/CER ↓
Emo2v ↑
SIM-I ↑
TCD ↓
OVRL ↑
UTMOS ↑
English
Emotional Dialogue
Concat.
2.93
90.18
0.659
0.452
3.807
3.462
Joint
0.46
65.22
0.912
0.463
3.708
3.752
TED-TTS
0.34
84.58
0.913
0.487
3.622
3.719
EDICT-Stage II
1.23
88.58
0.927
0.342
3.699
3.750
Observational Narrative
Concat.
0.43
84.62
0.652
0.469
3.820
3.525
Joint
0.12
58.68
0.939
0.508
3.818
4.197
Appendix
Table 17: MED-TTS results by language and text category. WER/CER denotes English WER or Chinese CER in percent. Emo2v is a segment-level emotion classification accuracy in percent; SIM-I measures inter-segment speaker similarity. TCD is reported as 10−3TCD , and OVRL denotes DNSMOS OVRL. Best and second-best scores within each language and category are bold and underlined .