How can we verify whose music contributed to an AI-generated output? This paper demonstrates how input-based attribution can provide verifiable evidence of which audio sources were used in a generation and whether they shaped the output. To do so, we condition the generation solely on audio without any text input, then trace the inputs behind each output, and establish their musical effect. In prompt adherence tests and controlled input swaps, the stems generated by our generator, MixAudio, follow the prompt audio in timbre and the context audio in harmony. Yet these outputs may still reproduce training data not supplied as inputs. We therefore audit memorization with our musical version identification model, musicDNA, and find few reproductions outside the input records. On human-judged cases within the flagged pool, it achieves higher precision and recall than the other tested memorization detectors. The two evaluations suggest that input records and output analysis provide complementary evidence for attribution, on which rights-holder reporting and compensation can draw as the AI music economy takes shape. Audio examples are available at https://neutune.github.io/attr2027demo/
Figures & tables
Figure 1 : Attribution by conditioning provenance. A single pass generates any subset of a track’s stems, conditioned exclusively on registered audio, with no text input: one prompt recording per stem and an optional shared context. The two rights panels on the left are the rights involved at generation. The dashed box is added when the context is mixed into the released master.
MERT (Style)
MFCC (Timbre)
Source
System
Matched
Random
Δ
Matched
Random
Δ
Training
MixAudio
0.550
0.014
0.535
0.914
0.554
0.360
ACE
0.549
0.121
0.428
0.838
0.602
0.236
Validation
MixAudio
0.516
0.000
0.516
0.911
0.550
0.361
ACE
0.534
0.123
0.411
0.831
0.602
0.229
MoisesDB
MixAudio
0.582
0.124
0.457
0.920
0.599
0.321
Table 1: Prompt adherence. Similarity between generated stems and prompts. Matched : each stem against the prompt it was generated from. Random : the same generations scored against another track’s prompt. Δ = Matched − Random. Bold: larger Δ per source.
MixAudio
ACE
Hypothesis
Metric
Base
Swap
Base
Swap
(a) Prompt replaced
Timbre moves to the substituted prompt
MFCC
0.63
0.89
0.69
0.85
Timbre leaves the original prompt
MFCC
0.92
0.64
0.88
0.70
Harmony stays with the original context
Chroma
0.20
0.22
0.10
0.06
(b) Context replaced
Table 2: Conditioning swaps. Each value is the similarity between a generated stem and the prompt or context marked in its row. Base : the base generation. Swap : the regeneration after the swap; the other conditioning input is held fixed. Averaged over the four dataset sources.
Figure 2 : How listeners judge a flagged stem, read top to bottom. Orange: Type I and II, the material unnamed memorizations; blue: named (A) or minor unnamed (D) memorization. Badges: n⋅ % of the band ⋅ % of the 994 queries, shown for the musicDNA band over both splits (59 stems).
Validation (unseen)
Training (registered)
Case
musicDNA
CLEWS
CLAP
MERT
musicDNA
CLEWS
CLAP
MERT
Flagged
25
25
25
25
34
35
31
24
Type I ⋅ Recording
0
0
0
0
1 (3%)
1 (3%)
0
0
Type II ⋅ Composition
0
0
0
0
0
0
0
0
A ⋅ Duplicate
9 (36%)
7 (28%)
1 (4%)
3 (12%)
19 (56%)
14 (40%)
0
7 (29%)
B ⋅ Unrelated
16 (64%)
12 (48%)
24 (96%)
22 (88%)
13 (38%)
16 (46%)
31 (100%)
17 (71%)
Table 3: Listener judgments on the flagged stems, per retriever and split, in the categories of Figure 2 . Percentages are shares of each column’s flagged stems.
Figure 3 : Detection performance of memorization metrics. One dot per human-judged stem, scored by all six detectors. Rows follow the categories of Figure 2 , in the same colors: colored dots are confirmed memorizations (Type I, A), grey dots are not (B, C). More similar is to the right, and the dashed line is each detector’s detection threshold. The Null row (grey band) marks chance-level similarity: the 5–95% range of scores for unseen validation stems. Vertical jitter within a row carries no meaning.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Tracks
Sections
Instances
Target stems
Contextless
Hours
Train
1922
2073
2073
5030
515
7.1
Val
168
2086
2086
4960
518
6.9
MoisesDB
112
1058
1058
2266
282
5.6
MUSDB18
150
1593
1593
2778
458
8.3
Appendix
Table 4: Evaluation set composition. One instance per section; target stems are the stems each system generates; contextless instances have no context mix. Hours: total section audio.
System
Prompt
Context
Stems
Why not a baseline
ACE-Step 1.5, Lego ( Gong et al., 2026 )
audio, per stem
✓
✓
baseline
MTMusicLDM ( Karchkhadze et al., 2026 )
text
✓
✓
no audio prompt
MSDM ( Mariani et al., 2024 )
–
inpainting
4 fixed
no audio prompt
YuE ( Yuan et al., 2026 )
audio
–
2 fixed
no context
MusicGen-Melody ( Copet et al., 2023 )
chroma
–
–
no context
Stable Audio Open ( Evans et al., 2025 )
audio-to-audio
–
–
no context
Appendix
Table 5 : Candidate baselines. Research systems with released weights (top) and commercial services (bottom). What each system takes as the prompt for the generated part and whether it takes existing audio as a context and produces stems (✓ yes, – no).
Hypothesis
Metric
Oracle
MixAudio
ACE
(a) Prompt time-scaled ( × 0.9 / 1.1)
Output tempo unchanged
Tempo accuracy
100%
95%
91%
Harmony stays with the context
Chroma Δ
≈0
−0.003
+0.003
Timbre unchanged
MFCC Δ
≈0
−0.005
−0.005
(b) Prompt pitch-shifted ( ±2 st)
Harmony stays with the context
Chroma Δ
≈0
−0.005
−0.025
Appendix
Table 6: Tempo and pitch transformations. One conditioning input is time-scaled or pitch-shifted by a known amount and the stem regenerated, while the other conditioning input is held fixed. Oracle is the value perfect following would yield.
System
Source
FAD ( ×10−3 ) ↓
Human ratings (MOS 1–5) ↑
Quality
Prompt
Context
MixAudio
Validation
0.596
2.81±0.29
3.36±0.32
3.45±0.39
MoisesDB
0.997
2.77±0.33
3.34±0.33
3.43±0.42
MUSDB18
1.010
2.80±0.27
3.18±0.31
3.20±0.30
All
0.868
2.80±0.17
3.29±0.19
3.36±0.22
ACE
Validation
0.935
2.36±0.25
2.73±0.34
2.56±0.44
Appendix
Table 8: FAD and listener ratings of generated stems, and listener ratings of the real stems of the same cases. FAD: lower is better. Ratings: mean opinion scores from the listening test (1–5) with 95% confidence intervals, higher is better. Quality, Prompt, and Context are sound quality, match to the prompt audio, and fit with the context.
Figure 4 : Listener ratings. Mean opinion scores over all 48 cases, the All rows of Table 8 , with their 95% confidence intervals.
Training data attribution (TDA) for music generation must answer two questions that copyright analysis requires, namely which training songs influence a generated output and along which musical aspects the influence operates. Existing methods reduce influence to a single scalar, without revealing which musical aspects are dominant in that influence. We propose ARIA, a framework that decomposes attribution along musical aspects (five for symbolic music, three for audio) and pairs the decomposition with reliability diagnostics computed from the segment-level score matrix. It measures within-group similarity among the top-K attributed tracks against random reference groups drawn from the training pool, and diagnoses the score matrix through its singular value decomposition and column statistics. On a symbolic-music model where attribution ground truth is available through counterfactual retraining, the reliability diagnostics rank four attribution methods identically to that ground truth. On an audio music generation model, ARIA reveals attribution behaviors that vary substantially across TDA methods, flags score matrices whose retrieved tracks are nearly identical across queries rather than reflecting per-query attribution, and characterizes embedding-similarity retrieval baselines by the musical aspect each encoder surfaces. Together, ARIA produces per-aspect attribution evidence aligned with the musical aspects considered under the idea-expression distinction in copyright analysis.
Changheon Han, Ashkan Panahi, Kıvanç Tatar
Department of Computer Science and Engineering, Chalmers University of Technology and University of Gothenburg Gothenburg, Sweden
The rapid expansion of text-to-music generative models challenges traditional paradigms of music creation and intellectual property. Plagiarism in this context is rarely an absolute mathematical binary, but an ambiguous threshold negotiated over harmonic structure, melodic contours, or overall perceived stylistic character. In this work, we test the transferability of state-of-the-art music version identification architectures from the human-to-human cover domain to the human-to-AI plagiarism setting. To evaluate this task, we introduce COPYCAT, a benchmark derived from real-world plagiarism cases and extended through generative re-synthesis and digital signal processing obfuscations, yielding 350,654 evaluation pairs. We show that scalar distance thresholding collapses under generative re-synthesis, while a supervised framework leveraging coordinate-wise embedding shifts recovers the dispersed plagiarism signal, raising overall F0.5 from 0.612 to 0.803.
Fotis Koutsikos, Ioannis Prokopiou, Spyridon Kantarelis +6
National Technical University of Athens, Athens, Greece · Athens University of Economics and Business, Athens, Greece · Orfium, Athens, Greece +1
We present ArtifactNet, a lightweight framework that detects AI-generated music by reframing the problem as forensic physics -- extracting and analyzing the physical artifacts that neural audio codecs inevitably imprint on generated audio. A bounded-mask UNet (ArtifactUNet, 3.6M parameters) extracts codec residuals from magnitude spectrograms, which are then decomposed via HPSS into 7-channel forensic features for classification by a compact CNN (0.4M parameters; 4.0M total). We introduce ArtifactBench, a multi-generator evaluation benchmark comprising 6,183 tracks (4,383 AI from 22 generators and 1,800 real from 6 diverse sources). Each track is tagged with bench_origin for fair zero-shot evaluation. On the unseen test partition (n=2,263), ArtifactNet achieves F1 = 0.9829 with FPR = 1.49%, compared to CLAM (F1 = 0.7576, FPR = 69.26%) and SpecTTTra (F1 = 0.7713, FPR = 19.43%) evaluated under identical conditions with published checkpoints. Codec-aware training (4-way WAV/MP3/AAC/Opus augmentation) further reduces cross-codec probability drift by 83% (Delta = 0.95 -> 0.16), resolving the primary codec-invariance failure mode. These results establish forensic physics -- direct extraction of codec-level artifacts -- as a more generalizable and parameter-efficient paradigm for AI music detection than representation learning, using 49x fewer parameters than CLAM and 4.8x fewer than SpecTTTra.
Heewon Oh
Intrect / MARTE Lab, Dongguk University, Seoul, South Korea