Instruction-based text-to-speech (TTS) offers control over voice characteristics and speech expression through interfaces including voice cloning and text-based voice design. Voice cloning reproduces a reference voice, whereas text-based voice design creates a voice from a natural-language description. However, neither interface directly enables users to modify the timbre of a given reference and synthesize speech with the modified voice. Meanwhile, utterance-level expressive instructions leave changes across individual text segments underspecified. We introduce \textbf{EDICT}, a framework that unifies global timbre editing and local expressive control by using an edited acoustic reference to anchor voice identity across segments. To enable synthesis with an instruction-edited voice, EDICT combines reference audio with structured timbre edits to generate an edited reference in codec-token space. This representation serves as a shared voice anchor for a frozen TTS backbone, allowing segment-specific natural-language instructions to guide expression. To accommodate instruction changes while supporting acoustic continuity, EDICT rebuilds the KV cache at each segment boundary, refreshing instruction conditioning while retaining bounded acoustic context from previously generated speech. Evaluations on our proposed TimbreEdit-Bench and IntraTTS-Bench demonstrate improved timbre editing and a favorable balance between local instruction adherence, speaker consistency, and transition quality. Audio demos are available.
Figures & tables
Figure 1: EDICT combines global voice editing with local delivery control. The trainable editor uses the source voice and global edit instruction to generate an edited codec reference. This reference conditions every segment of the frozen TTS backbone, whose KV cache is rebuilt at instruction boundaries using selected acoustic context. The segments form one continuous utterance.
Figure 2: Segment-boundary KV cache reconstruction. Each non-final <EOS> triggers a fresh prefill containing the next instruction and text, the shared edited reference, and acoustic inputs from the first S and most recent K generated positions. The upper-left panel shows replacement of the old cache; the upper-right panel shows causal attention after reconstruction. The lower timeline illustrates successive instruction switches within one codec stream.
Figure 3: Benchmark construction and evaluation. Left: TimbreEdit-Bench pairs source and target voices with editing instructions. Right: IntraTTS-Bench aligns delivery instructions with successive text segments.
System
Instruction
Content and quality
WavLM
ERes2Net
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-T ↑
SIM-R
Δ↑
SIM-T ↑
SIM-R
Δ↑
Instruction-conditioned systems
Qwen3-TTS-VD
Structured
4.5 / 0.3
3.861
3.946
0.701
0.871
−0.170
0.450
0.745
−0.295
Descriptive
3.0 / 0.3
3.830
3.965
0.676
0.832
−0.156
0.435
0.714
−0.279
Qwen3-TTS-Base
Structured
5.1 / 1.9
3.742
3.805
0.632
0.950
−0.318
0.431
0.845
−0.414
Descriptive
3.3 / 1.1
3.756
3.776
0.637
0.950
−0.313
0.434
0.849
−0.415
Table 1: Objective results on TimbreEdit-Bench. WER/CER reports English WER / Chinese CER in percent. SIM-T/R denotes target/source similarity; Δ=SIM-T−SIM-R . Best and second-best scores within each instruction form are bold and underlined , with CER and WER ranked separately; SIM-R is unranked. Target-reference voice conversion receives target audio and is evaluated separately on a same-text subset, with its best similarity and Δ scores in bold .
System
Instruction
Up Acc. ↑
Down Acc. ↑
Category Acc. ↑
Preservation ↑
Qwen3-TTS-VD
Structured
67.0
64.3
53.2
83.3
Descriptive
60.0
63.0
49.8
85.0
Qwen3-TTS-Base
Structured
62.3
54.5
46.4
45.7
Descriptive
66.0
54.5
50.0
45.6
CosyVoice 2
Structured
68.1
65.5
57.7
68.2
Descriptive
72.2
73.3
69.0
66.2
Table 2: Attribute-operation accuracy (%) on 500 TimbreEdit-Bench requests. Gemini compares source and generated audio without the target recording. Up/Down assesses ordinal direction, Category assesses categorical targets, and Preservation assesses unchanged attributes. Best and second-best scores within each instruction form are bold and underlined .
System
Control (%)
Consistency
Content and quality
SegAcc ↑
EmoAcc ↑
Trans ↑
SIM-I ↑
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
Qwen3-TTS-VD (Concat.)
92.44
92.66
60.68
0.658
4.5 / 0.9
3.266
3.939
Qwen3-TTS-VD (Joint)
66.80
66.37
70.66
0.922
3.8 / 1.3
3.794
3.952
Qwen3-TTS-VD (TED-TTS)
74.25
74.02
78.03
0.911
4.5 / 1.7
3.655
3.853
Qwen3-TTS-VD (EDICT-Stage II)
82.30
83.60
82.00
0.915
5.1 / 1.8
3.631
3.930
CosyVoice 2 (Concat.)
62.43
62.73
31.80
0.917
6.5 / 3.6
3.916
3.866
Table 3: Intra-utterance instruction control on IntraTTS-Bench. SegAcc, EmoAcc, and Trans measure instruction adherence, emotion realization, and transition quality in percent. SIM-I measures inter-segment speaker similarity; WER/CER reports English WER / Chinese CER in percent. Best and second-best scores within each backbone are bold and underlined , with CER and WER ranked separately.
Configuration
Control (%)
Content and quality
Consistency
SIM-T ↑
SegAcc ↑
Trans ↑
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-I ↑
WavLM
ERes2Net
Qwen3-TTS-VD (Joint)
60.19
70.07
3.6 / 1.2
3.339
3.869
0.918
0.683
0.395
Qwen3-TTS-VD (Concat.)
86.43
56.69
4.5 / 0.8
3.669
3.895
0.884
0.694
0.405
EDICT Stage I + Concat.
91.17
58.47
4.8 / 1.1
3.296
3.739
0.881
0.749
0.598
EDICT
81.62
83.45
5.1 / 1.9
3.795
3.921
0.917
0.788
0.609
Table 4: Joint timbre and delivery control with Qwen3-TTS-VD. Requests combine TimbreEdit-Bench voice edits with IntraTTS-Bench texts and delivery instructions. SIM-T compares the complete output with the target voice; SIM-I measures inter-segment speaker similarity. WER/CER reports English WER / Chinese CER in percent. Best and second-best scores are bold and underlined , with CER and WER ranked separately.
Figure 4: Subjective evaluation of timbre editing and local control. Bars show mean five-point ratings from 19 listeners; error bars are computed from 95% bootstrap confidence intervals over audio items. Each system has 10 items in (a) and 12 in (b). EDICT is highlighted in coral; all variants in (b) use Qwen3-TTS-VD. † Seed-VC receives target audio and is not rated for instruction-based edit accuracy.
Configuration
Instruction
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-T ↑
WavLM
ERes2Net
SFT (0.5M)
Structured
3.9 / 1.5
3.689
3.895
0.736
0.498
Descriptive
4.2 / 1.5
3.708
3.885
0.698
0.481
SFT (1M)
Structured
4.1 / 0.9
3.936
3.896
0.757
0.561
Descriptive
3.9 / 0.9
3.666
3.913
0.693
0.556
DPO ( M=4 )
Structured
5.1 / 1.5
3.885
3.962
0.761
0.552
Table 5: Timbre-editor training ablations on TimbreEdit-Bench. All post-training variants start from SFT (1M). M denotes the candidate count per condition; wDPO includes auxiliary SFT as defined in Equation 2 . SIM-T measures target-voice similarity; WER/CER reports English WER / Chinese CER in percent. Best and second-best scores within each instruction form are bold and underlined , with CER and WER ranked separately.
Configuration
Control (%)
Consistency
Content and quality
SegAcc ↑
EmoAcc ↑
Trans ↑
SIM-I ↑
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
Persistent KV cache
54.25
54.02
68.00
0.918
3.9 / 0.8
3.555
3.953
Truncated KV cache
58.05
56.70
64.94
0.892
4.4 / 1.1
3.278
3.905
EDICT-Stage II
82.30
83.60
82.00
0.915
5.1 / 1.8
3.631
3.930
w/o initial context
89.29
89.00
79.99
0.863
4.3 / 1.4
3.585
3.896
w/o recent context
81.15
82.26
55.32
0.875
4.2 / 1.1
3.548
3.830
Table 6: KV cache reconstruction ablations on IntraTTS-Bench. All variants use Qwen3-TTS-VD. The full configuration retains the first S=4 and most recent K=5 acoustic positions; the context ablations remove one of these groups. Persistent and truncated caches reuse old KV states, whereas EDICT recomputes them. Best and second-best scores are bold and underlined , with WER and CER ranked separately.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Control input
Control target
Scope
Speech content editing
Recording + revised text
Spoken content
Selected span
Timbre editing
Reference + edit instruction
Selected timbre attributes
Utterance
Voice cloning
Reference speech
Voice identity
Utterance
Voice design
Description / attribute specification
Voice characteristics
Utterance
Global expressive control
Utterance-level instruction
Emotion, rate, prosody, etc
Utterance
Local expressive control
Span-specific instructions
Emotion, rate, prosody, etc
Word / segment
Appendix
Table 7: Taxonomy of speech control tasks. Typical control inputs, control targets, and temporal scope are shown. Inputs are not exhaustive, and tasks may overlap.
Method
Zero-shot cloning
Voice design
Timbre modification
Intra-utterance control
Control formulation
VoxInstruct ( Zhou et al., 2024 )
✓
✓
✗
✓ a
Instruction-guided synthesis and word-level emphasis
Qwen3-TTS-Base ( Hu et al., 2026a )
✓
✗
✗
✗
Reference-based voice cloning
Qwen3-TTS-VoiceDesign ( Hu et al., 2026a )
✓
✓
✗
✗
Description-guided voice design
Spark-TTS ( Wang et al., 2025b )
✓
✓
✗
✗
Attribute-guided design and reference cloning
VoiceSculptor ( Hu et al., 2026b )
✓
✓
✗
✗
Voice design followed by cloning
CosyVoice 2 ( Du et al., 2024 )
✓
✗
✗
✓ b
Instructed cloning and phrase-level tags
Appendix
Table 8: Comparison of speech generation and editing methods. Capabilities refer to the named models in the cited settings. Superscripts specify restricted scope or additional training.
Attribute
Values
Perceptual description
Perceived gender
Masculine, Feminine
Perceived vocal masculinity or femininity.
Age impression
Child, Young adult, Mature
Perceived age conveyed by the voice.
Pitch register
Low, Mid, High
Typical pitch range of the voice.
Brightness
Dark, Neutral, Bright
Perceived prominence of high-frequency energy.
Vocal weight
Thin, Medium, Thick
Perceived fullness and body of the voice.
Resonance focus
Head, Balanced, Chest
Resonance perceived as head- or chest-focused.
Appendix
Table 9: Perceptual attribute schema for timbre editing. Perceived gender and age impression are categorical; the seven remaining attributes use ordinal levels 1–3 in the listed order.
Figure 5: Synthetic speech-bank statistics. (a) Retained recording durations in 0.5-second bins. (b,c) Perceived gender and age impression. (d) Distributions of the seven ordinal attributes; levels 1–3 follow the schema in Table 9 . Profile percentages are normalized within the training (5,800) and held-out (200) subsets. Young denotes young adult.
Figure 6: Edit coverage across one million training pairs. (a) Number of changed attributes per pair. (b,c) Source-to-target transitions for perceived gender and age impression. (d) Signed target-minus-source differences for ordinal attributes; Keep denotes no change. Heatmaps use a shared color scale and show percentages of all pairs. Each complete matrix in (b,c) and each column in (d) sums to 100% before rounding. Young denotes young adult.
Figure 7: IntraTTS-Bench composition. (a,b) Request counts by segment count and topic, split by language. (c,d) Segment lengths in bins of width two, normalized within each language. Chinese lengths count non-whitespace characters, including punctuation; English lengths count whitespace-separated words.
Criterion
Valid annotations
Agreement (%)
H1
H2
Gemini
H1–H2
H1–Gemini
H2–Gemini
SegAcc
197
198
200
90.4
83.8
85.9
EmoAcc
197
198
200
85.8
83.2
85.4
Trans
99
100
100
91.9
93.9
96.0
Appendix
Table 10: Human–Gemini agreement on intra-utterance control. Valid counts refer to segment-level judgments for SegAcc and EmoAcc, and boundary-level judgments for Trans. Agreement (%) uses jointly valid items for each pair.
Criterion
Valid annotations
Agreement (%)
H1
H2
Gemini
H1–H2
H1–Gemini
H2–Gemini
Up
100
100
100
81.0
79.0
76.0
Down
100
100
100
73.0
67.0
66.0
Category
100
100
100
82.0
84.0
78.0
Preservation
100
100
100
74.0
72.0
70.0
Appendix
Table 11: Human–Gemini agreement on timbre editing. Each annotator provides 100 valid judgments per criterion. Agreement (%) is computed from paired pass/fail labels.
Figure 8: Instructions shown to listening-test participants. Panel (a) gives general guidance on playback and rating; panels (b) and (c) describe the timbre-editing and intra-utterance control tasks, respectively. System identities are hidden.
Figure 9: Listening-test interfaces. Panel (a) collects naturalness, edit accuracy, and target timbre similarity ratings in sequence; the target reference is revealed only after edit accuracy is rated. Panel (b) collects naturalness, instruction following, and transition quality ratings for the complete utterance. All criteria use five-point scales.
Figure 10: Attention across segment transitions. Rows show zero-indexed Transformer layers 7 and 14; columns compare TED-TTS and EDICT on the same request, using separate rollouts. We average attention over all 16 heads in each layer. G denotes the global instruction, Ik/xk the segment instruction/text, H the preceding audio, S/R the initial/recent context, and c2 the new segment’s audio positions. Gray cells are masked or unavailable. All panels share a power color scale with γ=0.35 .
Figure 11: Attention during KV cache reconstruction. Each layer group shows replay with (a) old KV states and old conditions, (b) reconstructed KV states and new conditions, and (c) the full prefill followed by decoding. The dashed coral box marks retained acoustic inputs whose KV states are recomputed; solid lines separate prefill from decoding. Layers 7 and 14 are zero-indexed, with attention averaged over all 16 heads. The replay comparison changes both cache and conditions, so it does not isolate reconstruction alone.
Figure 12: Statistics of the ESD evaluation set. (a) Speaker-pair coverage. (b) Number of edited attributes per request. (c) Attribute-change distributions. In-bar numbers indicate case counts.
System
Instruction
Content and quality
WavLM
ERes2Net
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-T ↑
SIM-R
Δ↑
SIM-T ↑
SIM-R
Δ↑
(a) English
Instruction-conditioned systems
Qwen3-TTS-VD
Structured
5.3
4.219
3.487
0.510
0.714
−0.204
0.383
0.531
−0.148
Descriptive
6.1
4.162
3.400
0.483
0.709
−0.226
0.385
0.533
−0.148
Qwen3-TTS-Base
Structured
4.7
4.358
3.387
0.396
0.934
−0.538
0.301
0.770
−0.469
Appendix
Table 12: Objective results on the ESD evaluation set. WER/CER reports English WER or Mandarin CER. SIM-T/R denotes target/source similarity; Δ=SIM-T−SIM-R . Best and second-best scores within each language and instruction form are bold and underlined ; SIM-R is unranked. Target-reference VC is shown separately, with its best similarity and Δ scores in bold .
Configuration
Instruction
Content and quality
WavLM
ERes2Net
WER/CER (%) ↓
UTMOS ↑
DNSMOS ↑
SIM-T ↑
SIM-T ↑
(a) SFT data scale
SFT (0.5M)
Structured
3.9 / 1.5
3.689
3.895
0.736
0.498
Descriptive
4.2 / 1.5
3.708
3.885
0.698
0.481
SFT (1M)
Structured
4.1 / 0.9
3.936
3.896
0.757
0.561
Descriptive
3.9 / 0.9
3.666
3.913
0.693
0.556
Appendix
Table 13: Extended timbre-editor training ablations on TimbreEdit-Bench. SIM-T measures target-voice similarity; WER/CER reports English WER / Chinese CER in percent. Best and second-best scores within each instruction form are bold and underlined , with CER and WER ranked separately.
Figure 13: Sensitivity to the auxiliary SFT coefficient in wDPO on TimbreEdit-Bench. Solid and dashed curves denote structured and descriptive instructions, respectively. Shading marks the default λSFT=0.01 . The horizontal axis is linear up to 0.01 and logarithmic thereafter.
Initial context S
0
1
2
3
4
5
SIM-I ↑
0.872
0.885
0.902
0.899
0.912
0.915
EmoSim ↑
0.856
0.850
0.842
0.837
0.836
0.821
Appendix
Table 14: Effect of initial acoustic context on a 500-request IntraTTS-Bench subset. With Qwen3-TTS-VD and K=5 , we vary the number S of initial acoustic positions retained during reconstruction. SIM-I measures speaker consistency; EmoSim measures emotion-embedding similarity to independently synthesized reference segments. Bold marks the default configuration, rather than the highest score.
Recent context K
1
2
3
4
5
6
7
8
9
10
SIM-I ↑
0.876
0.895
0.903
0.911
0.912
0.912
0.915
0.917
0.916
0.912
EmoSim ↑
0.837
0.842
0.838
0.832
0.836
0.831
0.825
0.819
0.821
0.817
Appendix
Table 15: Effect of recent acoustic context on a 500-request IntraTTS-Bench subset. With Qwen3-TTS-VD and S=4 , we vary the number K of recent acoustic positions retained during reconstruction. SIM-I measures speaker consistency; EmoSim measures emotion-embedding similarity to independently synthesized reference segments. Bold marks the default configuration, rather than the highest score.
Language
System
WER/CER ↓
Emo2v ↑
SIM-I ↑
TCD ↓
OVRL ↑
UTMOS ↑
English
Concat.
1.18
88.33
0.667
0.467
3.799
3.550
Joint
0.30
59.47
0.929
0.487
3.754
4.040
TED-TTS
0.69
80.91
0.919
0.463
3.676
3.875
EDICT-Stage II
1.09
85.81
0.932
0.357
3.739
3.863
Chinese
Concat.
1.74
85.56
0.589
0.488
3.665
3.528
Joint
1.57
61.48
0.856
0.425
3.692
3.763
Appendix
Table 16: Intra-utterance emotion control on 500 MED-TTS examples. WER/CER denotes English WER or Chinese CER in percent. Emo2v is a segment-level emotion classification accuracy in percent; SIM-I measures inter-segment speaker similarity. TCD is reported as 10−3TCD , and OVRL denotes DNSMOS OVRL. Best and second-best scores within each language are bold and underlined .
Text category
System
WER/CER ↓
Emo2v ↑
SIM-I ↑
TCD ↓
OVRL ↑
UTMOS ↑
English
Emotional Dialogue
Concat.
2.93
90.18
0.659
0.452
3.807
3.462
Joint
0.46
65.22
0.912
0.463
3.708
3.752
TED-TTS
0.34
84.58
0.913
0.487
3.622
3.719
EDICT-Stage II
1.23
88.58
0.927
0.342
3.699
3.750
Observational Narrative
Concat.
0.43
84.62
0.652
0.469
3.820
3.525
Joint
0.12
58.68
0.939
0.508
3.818
4.197
Appendix
Table 17: MED-TTS results by language and text category. WER/CER denotes English WER or Chinese CER in percent. Emo2v is a segment-level emotion classification accuracy in percent; SIM-I measures inter-segment speaker similarity. TCD is reported as 10−3TCD , and OVRL denotes DNSMOS OVRL. Best and second-best scores within each language and category are bold and underlined .
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, parameters, or target region underspecified. We study a precise and explicit interface for speech editing: a transcript-grounded structural edit instruction with XML-style tags explicitly specifies typed operations and localizes them to transcript spans or boundaries. This semantic timeline avoids explicit timestamp alignment and provides an externally inspectable contract for compositional edits. We instantiate the interface in dots.tts.edit, an editor adapted from the continuous autoregressive dots.tts foundation model. Four representative speech-creation controls cover lexical content, affective expression, pitch and speaking-rate delivery, and temporal phrasing through text, emotion, prosody, and pause editing. Task-specific data pipelines construct operation- and scope-controlled pairs while retaining source-derived context outside each target region. We further introduce doteBench, a bilingual evaluation suite that measures precise instruction following, local preservation, and audio quality across the four controls and their composition. Experiments show leading overall instruction following and local preservation across its five editing categories, while audio quality remains comparable to existing open-source systems. Across three Seed-TTS-Eval shards, the model shows negligible differences from the base model in zero-shot TTS recognition error rate and speaker similarity.
Hankun Wang, Bohan Li, Shi Lian +7
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University · Xiaohongshu Inc.
Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consistency with surrounding unedited content. While prior work has shown that Supervised Fine-Tuning (SFT) enables TTS models to acquire functional editing capability, this approach remains fundamentally bottlenecked by imperfect paired editing data and coarse-grained optimization signals. To address these limitations, we propose CosyEdit2, a speech editing model built on a two-stage post-training framework that progresses from supervised editing initialization to editing-oriented Group Relative Policy Optimization (GRPO) over target-speech-free data. Extensive experiments demonstrate that CosyEdit2 not only substantially advances speech editing performance, but also unlocks better zero-shot TTS capability, revealing a deeper mutual relationship between the two tasks. Audio samples are available at https://cjy1018.github.io/CosyEdit2.
Junyang Chen, Yuhang Jia, Hui Wang +3
College of Computer Science, Nankai University · College of Artificial Intelligence, Nankai University
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
Hongyao Deng, Wenhao Guan, Xuetao Lin +4
School of Informatics, Xiamen University, China · School of Electronic Science and Engineering, Xiamen University, China