cs.CLSep 24, 2026

Parts-of-Speech as Emergent Categories in SAE Latent Space

Authors: Alessandro Bondielli, Lucia Passaro, Serena Auriemma, Alessandro Lenci

Organizations: CoLingLab, Department of Philology, Literature and Linguistics, University of Pisa · Department of Computer Science, University of Pisa

Abstract

Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

    Aug 11, 2026Nikolai Bolik, Lennart Stöpler, Artur AndrzejakSemantic RepresentationsModel Activations

  2. Persistent Sparse Autoencoders: Learning Feature-Specific Timescales in Language Model Representations

    Jul 19, 2026Haoyan Luo, Mateo Espinosa Zarlenga, Mateja JamnikImproving Sparse AutoencodersFuture Latent Representations

  3. Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

    Aug 22, 2025David Chanin, Adrià Garriga-AlonsoImproving Sparse AutoencodersModel Activations