Synthetic speech attribution aims to identify the generative system responsible for a speech signal, but current approaches typically rely on black-box neural networks that provide limited insight into their decisions. This work investigates prototype-based networks as an interpretable alternative, where predictions are grounded in comparisons with representative training examples. We adapt ProtoPNet to spectrogram-based speech representations and evaluate the proposed framework on the MLAAD dataset under closed-set, cross-lingual, and open-set conditions. Experiments show that prototype-based reasoning achieves competitive or improved attribution performance compared with the baseline while enabling example-based explanations. These results highlight that interpretability and performance can be jointly achieved in synthetic speech attribution through prototype-based modeling.
Figures & tables
Fig. 1 : Overview of the proposed prototype-based framework for interpretable synthetic speech attribution.
Training by
Model
BA ↑
Macro F1 ↑
TTS Instances
ResNet18
86.25%
0.8504
ProtoRN18 L4P3
87.48%
0.8674
ProtoRN18 L4P6
88.06%
0.8801
TTS Families
ResNet18
99.06%
0.9905
ProtoRN18 L4P3
99.27%
0.9928
ProtoRN18 L4P6
99.65%
0.9964
TABLE I : Comparison between ResNet18 and ProtoRN18 on the model and architecture classification tasks.
BA ↑
Macro F1 ↑
ProtoRN18 L3P1
98.93%
0.9893
ProtoRN18 L3P3
98.97%
0.9893
ProtoRN18 L3P6
98.83%
0.9902
ProtoRN18 L4P1
98.86%
0.9887
ProtoRN18 L4P3
99.27%
0.9928
ProtoRN18 L4P6
99.65%
0.9964
TABLE II : Ablation study on ProtoRN18 configurations.
Mel region
Prototype center
# Source classes
Low-frequency
<1000 Hz
1 (2.27%)
Mid-frequency
1000–3000 Hz
5 (11.36%)
High-frequency
>3000 Hz
38 (86.36%)
TABLE III : Distribution of learned prototypes activation center across frequency regions.
Fig. 2 : Prototype-based explanation for the test sample jane_eyre_11_f000168.wav generated by sesame_csm . The first panel shows the spatial similarity map between the input spectrogram and the selected prototype (higher values, stronger activation). The second localizes the maximum activation on the input spectrogram. The third shows the training spectrogram from which the prototype was grounded ( northandsouth_40_f000076.wav ) and its corresponding location.
Eval Setting
Model
Avg BA ↑
Avg BA Eng ↑
Drop ↓
All classes
ProtoRN18 L4P3
60.78%
99.46%
38.68 pp
ProtoRN18 L4P6
60.50%
99.71%
39.21 pp
Filtered
ProtoRN18 L4P3
88.63%
99.53%
10.89 pp
ProtoRN18 L4P6
89.07%
99.63%
10.56 pp
TABLE IV : Cross-lingual performance.
Language
TTS Families
BA ↑
BA Eng ↑
Drop ↓
Portuguese
5
96.66%
99.67%
3.01 pp
Spanish
9
89.00%
99.52%
10.52 pp
Italian
10
88.72%
99.48%
10.76 pp
Polish
7
87.98%
99.64%
11.66 pp
French
15
85.79%
99.55%
13.76 pp
Chinese
8
82.76%
99.45%
16.69 pp
TABLE V : ProtoRN18 L4P6 per-language performance.
Avg AUC ↑
Avg FPR@TPR95 ↓
Avg URR ↑
ProtoRN18 L4P3
89.84%
29.60%
68.83%
ProtoRN18 L4P6
83.76%
32.43%
67.16%
TABLE VI : Open-set performance. Results are reported in terms of AUC, FPR at 95% TPR and Unknown Rejection Rate (URR). URR threshold computed on the validation set.
AUC ↑
FPR@TPR95 ↓
URR ↑
Top 1
Top 2
Top 3
Index-TTS
93.58%
18.80%
79.60%
griffin_lim (33.05%)
FireRedTTS (20.15%)
Llasa (9.05%)
Kitten-TTS
88.36%
34.15%
64.95%
Spark-TTS-0.5B (43.00%)
Llasa (42.10%)
Chatterbox (12.95%)
MiniCPM-o-2.6
98.45%
4.70%
94.20%
FireRedTTS (13.10%)
orpheus-tts-0.1-finetune (11.50%)
MegaTTS3 (9.10%)
Kyutai-TTS
95.47%
17.10%
80.90%
sesame_csm (34.90%)
parler_tts (28.90%)
kokoro (6.30%)
Microsoft VibeVoice
93.22%
27.35%
70.00%
Spark-TTS-0.5B (30.30%)
Mars5 (27.65%)
MegaTTS3 (21.80%)
Veena
56.29%
97.30%
2.40%
orpheus-tts-0.1-finetune (99.70%)
Chatterbox (0.10%)
f5-tts (0.10%)
TABLE VII : Unknown class detection results and top-3 predicted source models by ProtoRN18 L4P3.
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, provided through dedicated techniques such as AudioSeal, or deployed by commercial platforms such as ElevenLabs. We identify a previously uncharacterized liability: when synthetic speech is watermarked and human speech is not, detectors trained alongside latch onto the watermark as a spurious "watermark => fake" shortcut. This single feature yields three coupled failures: generalization degradation (model performance deteriorates on unseen data), strip-to-evade (a watermarked fake escapes once unwatermarked), and mark-to-frame (watermarking a real voice flags it as fake). In a controlled white-box experiment, a watermark-trained detector shows all three (for example, mark-to-frame lifts Equal Error Rate from 16% to 75%). In a black-box test of a commercial API, we show that adding a watermark to real speech disguises it as fake. However, this shortcut is fixable: retraining with the watermark on both classes decorrelates it and restores clean behavior. We release experiment data as a paired clean-versus-watermarked corpus (WASP).
Localising artifacts in synthetic speech remains challenging, as most evaluation methods yield only global quality scores. This paper presents XSQ-AST, a framework that combines the SQ-AST speech quality model with WhisperX phoneme alignment and multiple saliency methods to produce temporally localised artifact diagnostics without model retraining. Saliency maps are projected onto continuous distributions via kernel density estimation and onto phoneme boundaries via phoneme-discretised saliency maps. A 40-participant listening test validated the framework across five perceptual dimensions. Attention Rollout, Attention Flow and an adapted GradCAM produced temporal distributions that correlated with listener highlights, with different methods best suited to different artifact types. An AUC-ROC analysis confirmed discrimination above chance.
Ben Heritage, Luca Resti, Mónica Villanueva Aylagas +3
AudioLab, School of Physics, Engineering and Technology, University of York, United Kingdom · Department of Computer Science, University of York, United Kingdom · SEED – Electronic Arts (EA), Sweden +1