cs.CVMay 9, 2026

When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation

Authors: Jörg Frochte

Organizations: Bochum University of Applied Sciences

Abstract

Raw cosine in the 768-dimensional output space of the Contrastive Style Descriptor (CSD) is now widely read as an absolute, calibrated style-fidelity score for text-to-image and style-imitation evaluation. We introduce the discrimination gap, a corpus-internal, prototype-free and threshold-free diagnostic that tests whether contrastive style cosines admit an absolute same-versus-different interpretation on a candidate artist corpus. On a 1799-artwork, 91-artist public-domain corpus, raw CSD cosine yields negative point-estimate gaps for 23/9123/91 artists at the pairwise level (2/912/91 robust under bootstrap) and for 15/9115/91 in the aggregated-pool scoring regime style-fidelity evaluations typically use. CSLS readout on the frozen backbone reduces the aggregated negative-gap count to 4/914/91; combined with positional-embedding interpolation to 336336 pixels it raises unsupervised pair-verification AUC from 0.8830.883 to 0.9050.905 across 2525 artist-disjoint splits. We refer to this diagnostic-driven readout protocol on the frozen backbone (CSLS as default, pos-interp 336336 as the stronger optional setting) as CSD+, not a new encoder.A cross-backbone check on CLIP-ViT-L/14, SigLIP-large and DINOv2-Large reproduces the same shared-tradition failure pattern, providing evidence that the residual reflects a shared limitation of the four backbones we tested rather than a CSD-specific artefact. Practical implication: before reporting CSD cosine as an absolute style-fidelity score, run the diagnostic on the candidate corpus; CSLS is the minimal correction when it fails.

Explore similar work

CardsList
  1. Evaluating Style-Personalized Text Generation: Challenges and Directions

    Aug 8, 2025Anubhav Jangra, Bahareh Sarrafzadeh, Silviu Cucerzan +2Stylistic FidelityWriting