cs.CLJun 30, 2026

Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law

Authors: Noah ScharrenbergChang Sun

Organizations: Department of Advanced Computing Sciences, Maastricht University, Maastricht, The Netherlands. · Contractuo, Eindhoven, The Netherlands.

Abstract

Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards focus narrowly on verbatim memorisation. EU copyright doctrine applies a broader standards: substantial similarity, which extends to stylistic choices, narrative structure, and creative elaboration. This mismatch between what current methods detect and what the law protects leaves a significant compliance gap. We introduce PSALM, an LLM-as-a-judge framework that operationalises EU copyright doctrine through ten evaluators assessing computational overlap, stylistic dimensions (writing style, narrative voice), content dimensions (character, plot, scene, world building), and statutory exceptions (parody, pastiche, quotation, scènes à faire). Applying PSALM to Llama~3.2 models fine-tuned on translated historical Dutch literary works, we find that: 1) instruction-tuned models exhibit non-trivial baseline stylistic similarity prior to corpus exposure; 2) fine-tuning induces systematic stylistic appropriation across all infringement-relevant dimensions, extending beyond verbatim memorisation to abstract narrative patterns; 3) Negative Preference Optimisation unlearning substantially reduces similarity but leaves detectable residual stylistic patterns. These findings indicate that safeguards targeting literal copying alone are insufficient to mitigate broader copyright risks. PSALM provides infrastructure for auditable, legally informed compliance evaluation, though the relationship between automated similarity scores and infringement determinations requires validation by legal experts. This work bridges qualitative legal standards and quantitative technical measurement, exposing fundamental tensions between generative AI and EU intellectual property law.

Explore similar work

CardsList
  1. One-shot Style Transfer LLM log-probabilities for Authorship Attribution and Verification

    Oct 15, 2025Pablo Miralles-González, Javier Huertas-Tato, Alejandro Martín +1AuthorshipStylometric

  2. CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs

    Sep 1, 2026Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi +3RightsClone