How can we verify whose music contributed to an AI-generated output? This paper demonstrates how input-based attribution can provide verifiable evidence of which audio sources were used in a generation and whether they shaped the output. To do so, we condition the generation solely on audio without any text input, then trace the inputs behind each output, and establish their musical effect. In prompt adherence tests and controlled input swaps, the stems generated by our generator, MixAudio, follow the prompt audio in timbre and the context audio in harmony. Yet these outputs may still reproduce training data not supplied as inputs. We therefore audit memorization with our musical version identification model, musicDNA, and find few reproductions outside the input records. On human-judged cases within the flagged pool, it achieves higher precision and recall than the other tested memorization detectors. The two evaluations suggest that input records and output analysis provide complementary evidence for attribution, on which rights-holder reporting and compensation can draw as the AI music economy takes shape. Audio examples are available at https://neutune.github.io/attr2027demo/
Figures & tables
Figure 1 : Attribution by conditioning provenance. A single pass generates any subset of a track’s stems, conditioned exclusively on registered audio, with no text input: one prompt recording per stem and an optional shared context. The two rights panels on the left are the rights involved at generation. The dashed box is added when the context is mixed into the released master.
MERT (Style)
MFCC (Timbre)
Source
System
Matched
Random
Δ
Matched
Random
Δ
Training
MixAudio
0.550
0.014
0.535
0.914
0.554
0.360
ACE
0.549
0.121
0.428
0.838
0.602
0.236
Validation
MixAudio
0.516
0.000
0.516
0.911
0.550
0.361
ACE
0.534
0.123
0.411
0.831
0.602
0.229
MoisesDB
MixAudio
0.582
0.124
0.457
0.920
0.599
0.321
Table 1: Prompt adherence. Similarity between generated stems and prompts. Matched : each stem against the prompt it was generated from. Random : the same generations scored against another track’s prompt. Δ = Matched − Random. Bold: larger Δ per source.
MixAudio
ACE
Hypothesis
Metric
Base
Swap
Base
Swap
(a) Prompt replaced
Timbre moves to the substituted prompt
MFCC
0.63
0.89
0.69
0.85
Timbre leaves the original prompt
MFCC
0.92
0.64
0.88
0.70
Harmony stays with the original context
Chroma
0.20
0.22
0.10
0.06
(b) Context replaced
Table 2: Conditioning swaps. Each value is the similarity between a generated stem and the prompt or context marked in its row. Base : the base generation. Swap : the regeneration after the swap; the other conditioning input is held fixed. Averaged over the four dataset sources.
Figure 2 : How listeners judge a flagged stem, read top to bottom. Orange: Type I and II, the material unnamed memorizations; blue: named (A) or minor unnamed (D) memorization. Badges: n⋅ % of the band ⋅ % of the 994 queries, shown for the musicDNA band over both splits (59 stems).
Validation (unseen)
Training (registered)
Case
musicDNA
CLEWS
CLAP
MERT
musicDNA
CLEWS
CLAP
MERT
Flagged
25
25
25
25
34
35
31
24
Type I ⋅ Recording
0
0
0
0
1 (3%)
1 (3%)
0
0
Type II ⋅ Composition
0
0
0
0
0
0
0
0
A ⋅ Duplicate
9 (36%)
7 (28%)
1 (4%)
3 (12%)
19 (56%)
14 (40%)
0
7 (29%)
B ⋅ Unrelated
16 (64%)
12 (48%)
24 (96%)
22 (88%)
13 (38%)
16 (46%)
31 (100%)
17 (71%)
Table 3: Listener judgments on the flagged stems, per retriever and split, in the categories of Figure 2 . Percentages are shares of each column’s flagged stems.
Figure 3 : Detection performance of memorization metrics. One dot per human-judged stem, scored by all six detectors. Rows follow the categories of Figure 2 , in the same colors: colored dots are confirmed memorizations (Type I, A), grey dots are not (B, C). More similar is to the right, and the dashed line is each detector’s detection threshold. The Null row (grey band) marks chance-level similarity: the 5–95% range of scores for unseen validation stems. Vertical jitter within a row carries no meaning.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Tracks
Sections
Instances
Target stems
Contextless
Hours
Train
1922
2073
2073
5030
515
7.1
Val
168
2086
2086
4960
518
6.9
MoisesDB
112
1058
1058
2266
282
5.6
MUSDB18
150
1593
1593
2778
458
8.3
Appendix
Table 4: Evaluation set composition. One instance per section; target stems are the stems each system generates; contextless instances have no context mix. Hours: total section audio.
System
Prompt
Context
Stems
Why not a baseline
ACE-Step 1.5, Lego ( Gong et al., 2026 )
audio, per stem
✓
✓
baseline
MTMusicLDM ( Karchkhadze et al., 2026 )
text
✓
✓
no audio prompt
MSDM ( Mariani et al., 2024 )
–
inpainting
4 fixed
no audio prompt
YuE ( Yuan et al., 2026 )
audio
–
2 fixed
no context
MusicGen-Melody ( Copet et al., 2023 )
chroma
–
–
no context
Stable Audio Open ( Evans et al., 2025 )
audio-to-audio
–
–
no context
Appendix
Table 5 : Candidate baselines. Research systems with released weights (top) and commercial services (bottom). What each system takes as the prompt for the generated part and whether it takes existing audio as a context and produces stems (✓ yes, – no).
Hypothesis
Metric
Oracle
MixAudio
ACE
(a) Prompt time-scaled ( × 0.9 / 1.1)
Output tempo unchanged
Tempo accuracy
100%
95%
91%
Harmony stays with the context
Chroma Δ
≈0
−0.003
+0.003
Timbre unchanged
MFCC Δ
≈0
−0.005
−0.005
(b) Prompt pitch-shifted ( ±2 st)
Harmony stays with the context
Chroma Δ
≈0
−0.005
−0.025
Appendix
Table 6: Tempo and pitch transformations. One conditioning input is time-scaled or pitch-shifted by a known amount and the stem regenerated, while the other conditioning input is held fixed. Oracle is the value perfect following would yield.
System
Source
FAD ( ×10−3 ) ↓
Human ratings (MOS 1–5) ↑
Quality
Prompt
Context
MixAudio
Validation
0.596
2.81±0.29
3.36±0.32
3.45±0.39
MoisesDB
0.997
2.77±0.33
3.34±0.33
3.43±0.42
MUSDB18
1.010
2.80±0.27
3.18±0.31
3.20±0.30
All
0.868
2.80±0.17
3.29±0.19
3.36±0.22
ACE
Validation
0.935
2.36±0.25
2.73±0.34
2.56±0.44
Appendix
Table 8: FAD and listener ratings of generated stems, and listener ratings of the real stems of the same cases. FAD: lower is better. Ratings: mean opinion scores from the listening test (1–5) with 95% confidence intervals, higher is better. Quality, Prompt, and Context are sound quality, match to the prompt audio, and fit with the context.
Figure 4 : Listener ratings. Mean opinion scores over all 48 cases, the All rows of Table 8 , with their 95% confidence intervals.