cs.AIOct 5, 2026

A Testable Theory of Atomic Features

Authors: Kenny Peng, Jon Kleinberg, Nikhil Garg

Organizations: Cornell University

Abstract

We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This "recovery principle" yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different data share features prevalent in both, and sufficiently large SAEs recover both parent and child features. In contrast to conventional wisdom that SAE features are unstable and "split" as size increases, we find that these predictions hold on SAEs of sizes ranging from 512 to 131,072 trained on two large embedding models. From a theoretical perspective, our results suggest the promise of a scientific theory of representations based on atomic features. Practically, our results suggest the promise of scaling SAEs.

Figures & tables

Appendix figures & tables46 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

    Aug 22, 2025David Chanin, Adrià Garriga-AlonsoImproving Sparse AutoencodersModel Activations

  2. Persistent Sparse Autoencoders: Learning Feature-Specific Timescales in Language Model Representations

    Jul 19, 2026Haoyan Luo, Mateo Espinosa Zarlenga, Mateja JamnikImproving Sparse AutoencodersFuture Latent Representations