cs.CLSep 23, 2026

Contrastive Learning for Authorship Verification

Authors: Peter Kirby

Abstract

Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors of model performance. Based on these considerations, we develop a ModernBERT Bi-Encoder model that achieves 98.4% accuracy on the PAN21 authorship verification task.

Explore similar work

CardsList
  1. One-shot Style Transfer LLM log-probabilities for Authorship Attribution and Verification

    Oct 15, 2025Pablo Miralles-González, Javier Huertas-Tato, Alejandro Martín +1AuthorshipStylometric

  2. Where Does Authorship Signal Emerge in Encoder-Based Language Models?

    May 19, 2026Francis Kulumba, Guillaume Vimont, Laurent Romary +1AuthorshipEncoding Models