cs.CLMay 11, 2026

Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings

Authors: Benjamin IcardLila SaineroAlice BretonEvangelia ZveJean-Gabriel Ganascia

Organizations: LIP6, Sorbonne University, CNRS, France · Infopro Digital, France

Abstract

Large language models (LLMs) can convincingly imitate human writing styles, yet it remains unclear how much stylistic information is encoded in embeddings from any language model and retained after LLM rewriting. We investigate these questions in French, using a controlled literary dataset to quantify the effect of stylistic variation via changes in embedding dispersion. We observe that embeddings reliably capture authorial stylistic features and that these signals persist after rewriting, while also exhibiting LLM-specific patterns. These analytical results offer promising directions for authorship imitation detection in the era of language models.

Explore similar work

CardsList
  1. One-shot Style Transfer LLM log-probabilities for Authorship Attribution and Verification

    Oct 15, 2025Pablo Miralles-González, Javier Huertas-Tato, Alejandro Martín +1AuthorshipStylometric