cs.LGJun 13, 2026

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders

Authors: Silen NaihinLev Stambler

Organizations: Experiential Labs · Tear Labs

Abstract

Sparse autoencoders (SAEs) detect features via inner product, so a feature's activation scales with both its directional alignment and the input's norm. Features that fire on token norm therefore claim dictionary slots regardless of content alignment. This matters because sublayer normalization has already discarded the magnitude the score measures, so the encoder detects a quantity the model does not read. We replace the score with a learned blend of cosine similarity and input magnitude, letting the optimizer choose how much norm to use; a per-feature extension lets each feature decide independently. In both regimes, training is free to recover inner product but never does, with no feature ever choosing more than half-magnitude dependence. At matched reconstruction, the cosine encoder learns features that align with human-recognizable concepts far more often than standard, filling dictionary slots that inner product wastes on norm detectors. Loss reweighting that equalizes gradients barely closes the gap, confirming forward-pass score geometry as the lever. The advantage is not universal across tasks or depths, but we believe cosine scoring should be the default for dictionary learning on normalized representations.

Explore similar work

Jul 2, 2026cs.LG

Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability

Sparse autoencoders (SAEs) decompose internal activations of neural networks into sparse linear combinations of learned features by fitting an overcomplete dictionary WRm×n\mathbf{W}\in\mathbb{R}^{m\times n} with m<nm<n, and inferring a sparse code xRn\mathbf{x}\in\mathbb{R}^n from hWx\mathbf{h}\approx\mathbf{W}\mathbf{x}. This inference problem closely resembles the canonical setup of compressed sensing, but dense decoders requires O(mn)O(mn) learned values, which becomes costly at large feature counts. We introduce Expander SAEs: TopK SAEs whose decoder and tied encoder are supported on a left-dd-regular expander mask with dmd\ll m, learning only dndn decoder values while keeping the sparse-coding problem (m,n,k)(m,n,k) fixed. The same structure reduces storage and turns the matching-pursuit correlation step Wr\mathbf{W}^\top \mathbf{r} in OMP into an O(dn)O(dn) gather-and-reduce operation. Our experiments show that across Pythia-70M/160M, Qwen2.5-3B, and Llama-3.2-1B residual-stream activations, varying dd traces a consistent storage--fidelity frontier, and that at the most compressed modern-LM setting, Qwen2.5-3B with d=7d=7 uses 293×293\times fewer learned decoder values than the full dense decoder while retaining 8484% of dense CE-loss recovered. Control experiments show that the improved storage--fidelity tradeoff is driven by sparse, diverse decoder support structure rather than by fewer learned decoder values, and that when sparse and dense decoders are compared at matched parameter count, part of the remaining gap comes from encoder amortisation. On the theoretical side, we show that expansion and column flatness are sufficient for identifiability of noiseless kk-sparse codes, and we derive complementary sufficient conditions under which OMP recovers the support exactly.
Rodrigo Mendoza-Smith
Jun 26, 2026cs.CL

VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather than directly connected to the Transformer's token vocabulary. We introduce Vocabulary-Aligned Sparse Autoencoder (VASAE), a method that trains SAE features under vocabulary-aligned anchoring and assigns each feature an intrinsic token name: the token string whose embedding is nearest to that feature. Without reducing reconstruction quality compared with a standard SAE, VASAE produces dictionaries with vocabulary-aligned features. Using a 0.8 cutoff on the nearest-token alignment score, dictionaries trained on GPT-2-small post-residual streams align about 90% of features in layers 0--10. In Llama-3.1-8B, representative shallow and middle-layer dictionaries contain strongly aligned features, including 92.8% in the shallow layer, while the representative final-layer dictionary shows limited alignment. After subtracting the sentence-level mean sparse code, case studies show that many remaining intrinsic token names are relevant to nearby input tokens. These results suggest that vocabulary-aligned anchoring can connect learned features to intrinsic token names during training, complementing post hoc interpretation of learned dictionaries.
Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah +1
Sep 9, 2026cs.LG

A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram

Sparse autoencoders (SAEs) are increasingly used to recover interpretable features from neural-network activations, yet systematic feature co-occurrence can cause distinct features to be absorbed or merged. The MAIS-O43 open problem proposes a controlled experiment to characterize when recovery of a true synthetic dictionary gives way to feature merging as the nesting fraction γγ, sparsity penalty λλ, and dictionary size MM vary. We implement the specified protocol and evaluate 200 independently initialized fits across ten of the 165 grid cells. We observe zero full-dictionary recoveries and zero merges. Instead, every run converges to a reproducible diffuse phase: reconstruction is nearly perfect, but learned atoms typically remain far from the true features (median best cosine 0.5-0.7 against a 0.95 recovery criterion) and learned codes are an order of magnitude denser than the ground truth. This behavior persists under robustness checks and across the full 165-cell grid using standard minibatch Adam (3,300 additional fits). Since the global optimum of the exact sparse-coding objective is known to merge nested features in the two-feature case, these results suggest that trained SAEs need not reach the corresponding minima, and that the phase diagram of trained models may differ fundamentally from that of objective minimizers.
Alexis D. Plascencia