cs.CL · 2605.10606 Copy arXiv ID · May 11, 2026 Save Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings Authors: Benjamin Icard , Lila Sainero , Alice Breton , Evangelia Zve , Jean-Gabriel Ganascia
Organizations: LIP6, Sorbonne University, CNRS, France · Infopro Digital, France
Abstract Large language models (LLMs) can convincingly imitate human writing styles, yet it remains unclear how much stylistic information is encoded in embeddings from any language model and retained after LLM rewriting. We investigate these questions in French, using a controlled literary dataset to quantify the effect of stylistic variation via changes in embedding dispersion. We observe that embeddings reliably capture authorial stylistic features and that these signals persist after rewriting, while also exhibiting LLM-specific patterns. These analytical results offer promising directions for authorship imitation detection in the era of language models.
Explore similar work Oct 15, 2025 · Pablo Miralles-González, Javier Huertas-Tato, Alejandro Martín +1 Authorship Stylometric
Apr 24, 2026 · Tom van Nuenen Rewrite Linguistics
Apr 27, 2026 · Connor Baumler, Calvin Bao, Huy Nghiem +3 Stylistic Fidelity Ai-Assisted Writing
Oct 15, 2025 · cs.CL J/K move · Enter open · S save
Pablo Miralles-González, Javier Huertas-Tato, Alejandro Martín, David Camacho
Department of Computer Systems Technical University of Madrid Madrid
Computational stylometry studies writing style through quantitative textual patterns, enabling applications such as authorship attribution, identity linking, and plagiarism detection. Despite the relevance of language modeling to these tasks, the pre-training of modern large language models (LLMs) has been underutilized in authorship attribution and verification. We introduce an unsupervised framework that uses the log-probabilities of an LLM to measure style transferability between two texts. This framework takes advantage of the extensive Causal Language Modeling (CLM) pre-training, one-shot capabilities and scale of LLMs, avoiding explicit supervision. Our methods substantially outperform prompting-based unsupervised baselines in authorship verification at similar model sizes, and is competitive with or improves contrastive baselines in most settings with sufficient model scale. We further observe strong performance across non-English languages. The effectiveness of the proposed framework improves consistently with increasing model scale. In the case of authorship verification, we propose an additional mechanism that increases test-time computation to improve accuracy; enabling flexible trade-offs between computational cost and task performance.