Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion editing. To answer this question, we probe the editability of pretrained flow-matching and hybrid TTS models by constructing a controlled test set and systematically diagnosing editing effects along the generative trajectory. Our analysis reveals that pretrained TTS models are substantially editable in emotion, but such editability is architecture- and trajectory-dependent and can be disrupted by early flow-matching steps, while cross-speaker emotion transport carries additional acoustic attributes beyond emotion. To address these limitations, we propose SEmoEdit, the first training-free framework that formulates emotion editing as dynamic velocity transport between source and target emotions, enabling robust, flow-based speech emotion editing directly within pretrained TTS models. SEmoEdit unifies three core operations: emotion replacement, emotion erasure, and continuous emotion interpolation, requiring neither parameter updates nor task-specific optimization. To systematically evaluate these capabilities, we introduce SEmoEditBench, a dataset comprising 600 editing cases, and conduct extensive experiments across state-of-the-art (SOTA) models and backbones. Our results show that SEmoEdit is highly effective and broadly applicable, outperforming existing training-based and activation-steering methods. Ultimately, this work reveals that pretrained speech flows possess rich, latent emotion-editing capabilities, providing useful guidance for real applications. Code, benchmark, and Audio samples are available at https://github.com/imxtx/SEmoEdit.
Figures & tables
Figure 1: Target emotion2vec ( Ma et al., 2024 ) probability, change in transcription error ( Δ WER, in percentage points; pp), and change in speaker similarity ( Δ S-SIM) before and after editing.
Figure 2: The first two columns show onset drift (top) and WER change from the source in percentage points (bottom) versus retained emotion gain relative to full-step editing. The green region denotes ≥ 90% gain retention. The third column demonstrates the onset drift problem.
Figure 3: Framework of SEmoEdit for dynamic speech emotion editing.
Category
Method
Backbone
Metrics
Emotion replacement
TEP ↑
SES ↑
E-SIM ↑
DES ↑
Δ WER ↓
S-SIM ↑
UTMOS ↑
Training- based
Step-Audio-EditX
–
0.222
0.445
0.521
0.550
-0.048
0.567
3.091
dots.tts.edit
–
0.434
0.600
0.669
0.713
-0.035
0.272
2.381
Auk
–
0.259
0.372
0.539
0.541
0.073
0.631
2.639
Activation Steering
CoCoEmo
CosyVoice 2
0.082
0.192
0.446
0.399
-0.015
0.715
3.061
IndexTTS2
0.035
0.127
0.402
0.269
-0.024
0.776
2.659
Table 2: Main results on the same-dataset same-speaker setting of SEmoEditBench. The top three results for each metric are highlighted in bold with dark, medium, and light pink backgrounds representing the first , second , and third best performances, respectively.
Emotion replacement
Emotion erasure
Intensity control
Category
Method
Backbone
SS-MOS ↑
ES-MOS ↑
N-MOS ↑
SS-MOS ↑
ES-MOS ↑
N-MOS ↑
EIC-MOS ↑
Training- based
Step-Audio-EditX
–
3.67
2.88
3.75
4.20
2.35
4.00
–
Auk
–
3.38
2.80
3.52
4.00
2.95
4.05
–
dots.tts.edit
–
3.52
3.27
3.77
3.55
3.15
3.65
–
Activation Steering
CoCoEmo
IndexTTS2
4.42
1.88
3.98
–
–
–
1.65
EmoSteer-TTS
IndexTTS2
4.35
1.75
4.12
4.50
1.80
4.25
1.30
Table 3: Subjective evaluation results. The first , second , and third place are highlighted.
Figure 4: Left: SEmoEdit ID/OOD generalization across backbones. Right: Onset drift over 600 cases; lower mean and median absolute drift are better, while a higher fraction within 50 ms is better.
Figure 5: Sensitivity to the number of coupled noise samples per editing step.
Figure 6: Mean TEP trajectories before and after bridging.
Backbone
Variant
WER/CER ↓
S-SIM ↑
UTMOS ↑
F5-TTS
Raw
0.342
0.418
2.151
Bridged
0.205
0.431
2.411
CosyVoice 2
Raw
0.360
0.349
2.134
Bridged
0.201
0.457
2.780
IndexTTS2
Raw
0.823
0.379
1.918
Bridged
0.603
0.440
2.652
Table 4: Objective metrics before and after emotion bridging, averaged over 90 intensity editing cases.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: F5-TTS conditioning. The source-generated mel segment is linearly interpolated to the length of the target-generated mel segment.
Figure 8: CosyVoice 2 conditioning. Each branch uses its own reference mel and speaker embedding, together with the generated mel segment under that reference. The source branch’s generated mel is linearly interpolated to the target length, while the target branch remains unchanged. The two branches share a Gaussian noise sequence of length max(Ps,Pt) , which is right-aligned to each reference prefix.
Model
Emotion
Before (%, ↑ )
After (%, ↑ )
Δ (pp, ↑ )
F5-TTS
Happy
5.0456
64.8676
+59.8219
Angry
19.9294
77.8099
+57.8805
Sad
19.8356
60.8144
+40.9788
Surprise
0.2028
39.6302
+39.4274
Mean (80 cases)
11.2534
60.7805
+49.5272
CosyVoice 2
Happy
10.0841
70.5735
+60.4894
Appendix
Table 5: Target-emotion probability before and after editing. Values are percentages and Δ is in percentage points (pp). CosyVoice uses the aligned-source baseline.
Model
Emotion
Before (%, ↓ )
After (%, ↓ )
Δ (pp, ↓ )
F5-TTS
Happy
18.0595
15.3095
-2.7500
Angry
20.3591
9.4702
-10.8889
Sad
9.7123
8.9385
-0.7738
Surprise
12.4702
16.5575
+4.0873
Mean (80 cases)
15.1503
12.5689
-2.5813
CosyVoice 2
Happy
8.8095
13.8452
+5.0357
Appendix
Table 6: Per-case-averaged WER before and after editing. Positive changes indicate degradation. CosyVoice uses the aligned-source baseline; its whole-pipeline comparison is reported in the text.
Model
Emotion
Before ↑
After ↑
Δ↑
F5-TTS
Happy
0.6522
0.5966
-0.0556
Angry
0.6629
0.5695
-0.0934
Sad
0.6591
0.6186
-0.0405
Surprise
0.6878
0.5499
-0.1379
Mean (80 cases)
0.6655
0.5836
-0.0819
CosyVoice 2
Happy
0.6454
0.5716
-0.0738
Appendix
Table 7: Emotion-matched S-SIM defined in Eq. 12 . Before uses the neutral reference; after uses the target-emotion reference. CosyVoice before is the unedited generated speech before length alignment, so changes include alignment. Δ is an absolute cosine difference.
Model
Editing prefix
Emotion gain (pp)
Onset drift (ms)
Δ WER (pp)
F5-TTS
First 4 of 32 steps
0.015
0.38
+0.21
F5-TTS
All 32 steps
53.78
308.33
−0.92
CosyVoice 2
First 3 of 10 steps
3.43
22.21
+2.83
CosyVoice 2
First 6 of 10 steps
30.37
87.83
+88.88
CosyVoice 2
All 10 steps
56.09
245.42
+1.40
Appendix
Table 8: Early-stop editing results. Editing begins at the first editing step and terminates after the indicated prefix. Results are averaged over 80 cases and three editing-noise repeats.
Model
Editing start
Gain retained (%)
Onset drift (ms)
Δ WER (pp)
F5-TTS
Full-step
100.0
308.33
−0.92
Skip 4 steps
99.5
97.33
−4.95
CosyVoice 2
Full-step
100.0
245.42
+1.40
Skip 1 step
100.3
260.29
+0.67
Skip 3 steps
51.0
44.25
+55.75
Appendix
Table 9: Selected delayed starts, averaged over cases and editing-noise repeats. Gain retention is relative to full-step editing; onset drift and Δ WER are relative to the decoded, length-aligned source.
Figure 9: Complete delayed-start scans for F5-TTS and CosyVoice 2. Columns show emotion gain (pp), onset drift (ms), and Δ WER (pp), measured relative to the decoded, length-aligned source. Bands indicate pointwise 95% speaker-bootstrap intervals. The horizontal axis is the ODE step.
Model
Prefix
t
Emotion gain (pp)
Onset drift (ms)
Δ WER (pp)
F5-TTS
First 4 / 32
0.019
0.015
0.38
+0.21
F5-TTS
First 14 / 32
0.227
6.96
38.54
+1.31
F5-TTS
First 28 / 32
0.805
50.15
275.04
+6.03
F5-TTS
All 32 / 32
1.000
53.78
308.33
−0.92
CosyVoice 2
First 3 / 10
0.109
3.43
22.21
+2.83
CosyVoice 2
First 6 / 10
0.412
30.37
87.83
+88.88
Appendix
Table 10: Early-stop analysis. Editing starts at the first step and stops after the indicated prefix; values are means over 80 cases and three editing-noise repeats.
Figure 10: Early-stop scans for F5-TTS and CosyVoice 2. Columns report emotion gain, onset drift, and Δ WER. Bands show 95% speaker-bootstrap intervals; dashed lines indicate full-step means.
Model
Block
Only block
Skip block
Gain
Drift
Δ WER
Gain
Drift
Δ WER
F5-TTS
1
6.96
38.54
1.31
4.66
53.71
−0.97
2
1.17
14.17
0.60
36.00
205.38
13.88
3
0.72
0.50
0.25
40.27
253.75
8.10
4
0.23
2.62
−0.25
49.26
267.79
4.95
5
0.36
7.38
−0.24
50.15
275.04
6.03
Appendix
Table 11: Block ablations. Values are emotion gain (pp), onset drift (ms), and Δ WER (pp).
Figure 11: F5-TTS block and editing-strength ablations.
Figure 12: CosyVoice 2 block and editing-strength ablations.
Task
Setting
Corpus
Cases
Construction
Replacement
Same dataset, same speaker
ESD, IEMOCAP
120
80 ESD and 40 IEMOCAP cases; sources and targets balanced across Neutral, Happy, Sad, Angry, and Surprise; Chinese and English in ESD.
Replacement
Same dataset, cross speaker
ESD, IEMOCAP
120
Same source cases as the same-speaker split, but the target reference comes from a different speaker of the same corpus.
Replacement
Cross dataset, cross speaker
ESD, IEMOCAP
80
Sources from one corpus and target references from the other; all English.
Erasure
Same dataset, same speaker
ESD, IEMOCAP
56
32 ESD and 24 IEMOCAP cases; source emotions Happy, Sad, Angry, and Surprise, all mapped to Neutral.
Erasure
Same dataset, cross speaker
ESD, IEMOCAP
56
Same source cases, target reference from a different speaker of the same corpus.
Erasure
Cross dataset, cross speaker
ESD, IEMOCAP
40
Sources from one corpus and target references from the other; all English.
Appendix
Table 12: SEmoEditBench task composition. Each task is evaluated under same-dataset same-speaker, same-dataset cross-speaker, and cross-dataset cross-speaker settings.
Symbol
Meaning
xs , xe , xt
Source, edited, and paired-target waveforms, respectively
esrc , etgt
Source and target emotion labels
y
Manifest transcript (ground-truth text)
P(e∣x)
emotion2vec+ Large posterior probability for emotion e given waveform x
fE(x)
emotion2vec+ Large embedding of waveform x
A(x)
ASR transcript of waveform x produced by Whisper-large-v3
Appendix
Table 13: Notation used in the metric definitions.
Metric
Definition
Applicability
Editing success
Target emotion probability (TEP) ↑
P(etgt∣xe)
Replacement
Neutral probability (NP) ↑
P(Neutral∣xe)
Erasure
Source emotion suppression (SES) ↑
P(esrc∣xs)−P(esrc∣xe)
Replacement, erasure
Embedding-based effective intensity control (EIC-Emb) ↑
Mean over all pairs i<j of sgn(aj−ai)⋅min(1,max(0,∣aj−ai∣/τ)) , τ=1
Intensity control (paired target)
Emotion similarity (E-SIM) ↑
cos(fE(xe),fE(xt))
Splits with paired targets
Appendix
Table 14: Metrics used in SEmoEditBench. Upward and downward arrows indicate whether higher or lower values are preferred.
Category
Method
Backbone
Setting
TEP ↑
SES ↑
E-SIM ↑
DES ↑
Δ WER ↓
S-SIM ↑
UTMOS ↑
Training- based
Step-Audio-EditX
–
ID-SS
0.222
0.445
0.521
0.550
-0.048
0.567
3.091
ID-CS
0.196
0.420
0.519
0.543
-0.021
0.561
3.053
OOD
0.211
0.452
0.526
0.575
-0.026
0.519
3.142
Auk
–
ID-SS
0.259
0.372
0.539
0.541
0.073
0.631
2.639
ID-CS
0.259
0.372
0.539
0.541
0.073
0.631
2.670
OOD
0.312
0.353
0.586
0.580
0.103
0.580
2.753
Appendix
Table 15: Complete emotion-replacement results under both in-distribution settings and the out-of-distribution setting.
Category
Method
Backbone
Setting
NP ↑
SES ↑
E-SIM ↑
DES ↑
Δ WER ↓
S-SIM ↑
UTMOS ↑
Training- based
Step-Audio-EditX
–
ID-SS
0.235
0.388
0.566
0.655
-0.064
0.577
3.220
ID-CS
0.347
0.488
0.606
0.688
-0.028
0.586
3.133
OOD
0.269
0.465
0.746
0.818
-0.070
0.561
3.197
Auk
–
ID-SS
0.229
0.311
0.528
0.582
0.014
0.676
2.859
ID-CS
0.229
0.311
0.528
0.582
0.014
0.676
2.881
OOD
0.272
0.371
0.591
0.585
0.033
0.651
2.959
Appendix
Table 16: Complete emotion-erasure results under both in-distribution settings and the out-of-distribution setting. CoCoEmo is omitted because it does not support emotion erasure.
Category
Method
Backbone
Setting
EIC-Emb ↑
Δ WER α=1 ↓
S-SIM α=1 ↑
UTMOS α=1 ↑
Activation- steering
CoCoEmo
CosyVoice 2
ID-SS
0.041
0.044
0.621
3.234
ID-CS
0.023
0.006
0.648
3.206
OOD
0.057
0.044
0.595
2.990
CoCoEmo
IndexTTS2
ID-SS
0.017
0.333
0.703
2.294
ID-CS
0.007
0.378
0.699
2.344
OOD
0.026
0.356
0.688
2.314
Appendix
Table 17: Complete emotion-intensity-control results under both in-distribution settings and the out-of-distribution setting. Training-based baselines are omitted because they do not support continuous intensity control.
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
Hongyao Deng, Wenhao Guan, Xuetao Lin +4
School of Informatics, Xiamen University, China · School of Electronic Science and Engineering, Xiamen University, China
Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consistency with surrounding unedited content. While prior work has shown that Supervised Fine-Tuning (SFT) enables TTS models to acquire functional editing capability, this approach remains fundamentally bottlenecked by imperfect paired editing data and coarse-grained optimization signals. To address these limitations, we propose CosyEdit2, a speech editing model built on a two-stage post-training framework that progresses from supervised editing initialization to editing-oriented Group Relative Policy Optimization (GRPO) over target-speech-free data. Extensive experiments demonstrate that CosyEdit2 not only substantially advances speech editing performance, but also unlocks better zero-shot TTS capability, revealing a deeper mutual relationship between the two tasks. Audio samples are available at https://cjy1018.github.io/CosyEdit2.
Junyang Chen, Yuhang Jia, Hui Wang +3
College of Computer Science, Nankai University · College of Artificial Intelligence, Nankai University
Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing. To address these limitations, we propose CosyEdit, an end-to-end speech editing model adapted from CosyVoice through task-specific post-training and a complementary training paradigm, which internalizes text--speech alignment while ensuring high consistency between the speech before and after editing. Trained on only 250 hours of supervised data from our curated GigaEdit dataset, our 400M-parameter model achieves reliable speech editing performance. Extensive evaluations show that CosyEdit not only outperforms several billion-parameter language model baselines but also approaches state-of-the-art cascade systems. These results show that robust and efficient speech editing can be unlocked from a zero-shot TTS model through post-training, offering a cost-effective end-to-end solution for high-quality speech editing. Code and audio samples are available at https://cjy1018.github.io/CosyEditDemoPage/.
Junyang Chen, Yuhang Jia, Hui Wang +2
College of Computer Science, Nankai University, Tianjin, China