Towards AI-Generated Music Plagiarism Detection as a Version Identification Problem
Authors: Fotis Koutsikos, Ioannis Prokopiou, Spyridon Kantarelis, Vassilis Lyberatos, Pantelis Vikatos, Athanasios Aidinis, Themos Stafylakis, Athanasios Voulodimos, +1 more
Organizations: National Technical University of Athens, Athens, Greece · Athens University of Economics and Business, Athens, Greece · Orfium, Athens, Greece · Archimedes/Athena R.C., Athens, Greece
The rapid expansion of text-to-music generative models challenges traditional paradigms of music creation and intellectual property. Plagiarism in this context is rarely an absolute mathematical binary, but an ambiguous threshold negotiated over harmonic structure, melodic contours, or overall perceived stylistic character. In this work, we test the transferability of state-of-the-art music version identification architectures from the human-to-human cover domain to the human-to-AI plagiarism setting. To evaluate this task, we introduce COPYCAT, a benchmark derived from real-world plagiarism cases and extended through generative re-synthesis and digital signal processing obfuscations, yielding 350,654 evaluation pairs. We show that scalar distance thresholding collapses under generative re-synthesis, while a supervised framework leveraging coordinate-wise embedding shifts recovers the dispersed plagiarism signal, raising overall F0.5 from 0.612 to 0.803.
Figures & tables
Figure 1: Overview of the proposed framework.
COPYCAT
WEALY
CLEWS
XGBoost Hybrid Top-512
Category
Pairs
%
Prec.
Rec.
F 0.5
Prec.
Rec.
F 0.5
Prec.
Rec.
F 0.5
Human Plagiarism
3,757
3.3
74.2%
25.9%
54.0%
98.0%
30.6%
68.0%
97.3% ± 0.1%
86.5% ± 0.5%
94.9% ± 0.2%
Human Plagiarism + DSP
37,570
33.0
72.3%
24.8%
52.3%
97.7%
27.4%
64.5%
97.2% ± 0.1%
83.1% ± 0.4%
94.0% ± 0.2%
Original + DSP
12,190
10.7
51.8%
80.1%
55.7%
96.9%
95.2%
96.5%
87.2% ± 0.4%
99.3% ± 0.0%
89.4% ± 0.3%
AI Generation
5,472
4.8
42.4%
51.5%
43.9%
84.3%
15.0%
43.8%
75.1% ± 0.4%
56.5% ± 0.3%
70.5% ± 0.3%
AI + DSP
54,701
48.1
36.4%
34.3%
36.0%
79.0%
10.3%
33.9%
72.0% ± 0.5%
50.0% ± 0.3%
66.2% ± 0.4%
Table 1: Category-wise breakdown of the WEALY , CLEWS threshold baselines and the XGBoost Hybrid Top-512 classifier, each fitted globally and evaluated per category. XGBoost results are reported as mean ± 95% CI over 10 random seeds. Best F 0.5 per category in bold.
Figure 2: Origin-centered UMAP projection of CLEWS embeddings by transformation category (colors as in Fig. 3 ).
How can we verify whose music contributed to an AI-generated output? This paper demonstrates how input-based attribution can provide verifiable evidence of which audio sources were used in a generation and whether they shaped the output. To do so, we condition the generation solely on audio without any text input, then trace the inputs behind each output, and establish their musical effect. In prompt adherence tests and controlled input swaps, the stems generated by our generator, MixAudio, follow the prompt audio in timbre and the context audio in harmony. Yet these outputs may still reproduce training data not supplied as inputs. We therefore audit memorization with our musical version identification model, musicDNA, and find few reproductions outside the input records. On human-judged cases within the flagged pool, it achieves higher precision and recall than the other tested memorization detectors. The two evaluations suggest that input records and output analysis provide complementary evidence for attribution, on which rights-holder reporting and compensation can draw as the AI music economy takes shape. Audio examples are available at https://neutune.github.io/attr2027demo/
Taejun Kim, Wonil Kim, Jongmin Jung +10
Neutune, Seoul, South Korea · KAIST, Daejeon, South Korea
Recent advances in generative AI are raising ethical concerns regarding the originality of generated content and the potential replication of training data, with further implications for transparency, attribution, and intellectual property. In music, several computational approaches have been proposed to identify potential replication, using audio-based similarity metrics. Yet, their alignment with human judgments across distinct musical attributes remains underexplored. To address this gap, we conduct a perceptual experiment on music matches, defined as strongly similar musical excerpts. We focus on five musical attributes: melody, harmony, rhythm, voice, and timbre. We design a triplet-based forced-choice task comprising 300 cases, including plagiarism examples, cover songs, and AI-generated music. From this experiment, we introduce the MATCHA (Musical Attribute-based Triplet Comparison with Human Annotations) dataset: a collection of 1105 perceptual assessments of attribute-based music matches from 83 expert participants. Our findings reveal measurable agreement among participants in identifying matches across attributes. We further observe partial alignment between human judgments and computational similarity measures. Overall, this work underscores the importance of domain-specific and perceptually grounded evaluation frameworks for generative AI in creative practice.
Roser Batlle-Roca, Woosung Choi, Joan Serrà +5
Music Technology Group, Universitat Pompeu Fabra · Sony AI · Joint Research Centre, European Commission +1
We present a novel method for AI-generated music detection in scenarios where the models that generated the input samples are unknown to the detector (e.g., from a newly released service). Since 2023, there has been a multiplication of user-friendly AI-music generation services (e.g., Suno, Udio), along with regular updates and new features. There is thus a need to address synthetic content detection in an unsupervised way to adapt to this rapidly changing context. This angle has not been much studied in music yet. We propose to study two tasks. First, discriminating between real and synthetic music. This may be approached in a one-class manner, namely, using some baseline real music and trying to determine what falls outside. Second, zero-shot multi-class identification, which is more similar to an unsupervised clustering task on a mix of real and various AI-music generations, where the goal is to create coherent, high-purity clusters. We propose a combination of a previously proposed artifact-extraction method, on top of which we apply non-negative matrix factorization and simple classification and clustering methods. We achieve excellent performance on both tasks, showing that the proposed methods may be used to monitor large-scale catalogs that may receive AI-generated samples from various newly released generative models.