cs.CVJul 20, 2026

Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence

Authors: Katarzyna FilusSebastian Pokuciński

Organizations: Institute of Theoretical and Applied Informatics, Polish Academy of Sciences, Bałtycka 5, 44-100 Gliwice, Poland · Department of Applied Informatics, Silesian University of Technology, Akademicka 16, 44-100 Gliwice, Poland

Abstract

Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable features from neural representations. However, assessing their monosemanticity, and thus explanation quality, remains challenging. Existing metrics require external concept labels or depend on pretrained embedding models, making them sensitive to encoder's geometry. We introduce the Tversky Monosemanticity Score (TMS), a label-free metric that operationalizes monosemanticity as activation-set coherence of binarized SAE latents, and does not require external embedding encoders. We evaluate TMS on SAEs trained on features from pretrained vision and vision-language models (DINOv3, CLIP, BLIP2), two common SAE regimes (TopK, BatchTopK), multiple sparsity levels, and expansion factors. Our results show that TMS is less affected by encoder anisotropy than its embedding-based alternative, while remaining aligned with established monosemanticity indicators. TMS also reveals distinct SAE training dynamics across base models. Moreover, under encoder anisotropy, TMS provides a stronger indication of probe-based concept deletion effectiveness, while being competitive otherwise.

Explore similar work

CardsList
  1. The Rate-Distortion-Polysemanticity Tradeoff in SAEs

    May 14, 2026Tommaso Mencattini, Francesco Montagna, Francesco LocatelloSemantic RepresentationsMechanistic Interpretability