Sparse Autoencoders

Also known as SAE

Momentum

16 papers in the last four weeks, up 100% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 149

All topics
CardsList
  1. Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

    Jul 27, 2026Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty +2Sparse AutoencodersMechanistic Interpretability

  2. Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

    Jul 22, 2026Seonglae Cho, Zekun Wu, Kleyton Da Costa +3LLM InterpretabilitySparse Autoencoders

  3. Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence

    Jul 20, 2026Katarzyna Filus, Sebastian PokucińskiRepresentation DisentanglementSparse Autoencoders

  4. Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?

    Jul 19, 2026Aniket DeshpandeModel CompressionActivation Sparsity

  5. Persistent Sparse Autoencoders: Learning Feature-Specific Timescales in Language Model Representations

    Jul 19, 2026Haoyan Luo, Mateo Espinosa Zarlenga, Mateja JamnikSparse AutoencodersMechanistic Interpretability

  6. When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities

    Jul 9, 2026Weiduo Liao, Yunqiao Yang, Ying WeiVLM InterpretabilitySparse Autoencoders

  7. Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

    Jul 9, 2026Bendegúz Váradi, Zoltán KmettyTransformer InterpretabilitySparse Autoencoders

  8. A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models

    Jul 6, 2026Nima Eshraghi, Lovedeep Gondara, Yuqing Huang +5Language Model SteeringSparse Autoencoders

  9. Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models

    Jun 30, 2026Enrico Cassano, Riccardo Renzulli, Rayyan Ahmed +2Sparse AutoencodersDiffusion Model Unlearning

  10. Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

    Jun 30, 2026Jisung Park, Seohyeon Kang, Daeun Yoo +8Feature SuperpositionNeural Representation Geometry

  11. C2^{2}R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders

    Jun 29, 2026Haoran Jin, Xiting Wang, Shijie Ren +2LLM InterpretabilityRepresentation Disentanglement

  12. Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders

    Jun 29, 2026Chungpa Lee, Jihoon Kwon, Kyle Min +1Cross-Modal AlignmentSparse Autoencoders

  13. Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

    Jun 26, 2026Kevin Der, Harish Kamath, Ben ThompsonLLM InterpretabilitySparse Autoencoders

  14. VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

    Jun 26, 2026Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah +1LLM InterpretabilitySparse Autoencoders

  15. PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

    Jun 25, 2026Giosue Migliorini, Aristofanis Rontogiannis, Grigori Guitchounts +3Sparse AutoencodersMechanistic Interpretability

  16. Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

    Jun 25, 2026Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. HosseiniActivation SparsitySparse Autoencoders

  17. Discovering Millions of Interpretable Features with Sparse Autoencoders

    Jun 25, 2026XinYang He, Wei Wang, Bing Zhao +5LLM InterpretabilitySparse Autoencoders

  18. Steering Vision-Language Models with Joint Sparse Autoencoders

    Jun 24, 2026Huizhen Shu, Xuying Li, Hongxu Lin +2VLM InterpretabilitySparse Autoencoders

  19. Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

    Jun 23, 2026Jonas Klotz, Cassio F. Dantas, Pallavi Jain +2VLM InterpretabilitySparse Autoencoders

  20. Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?

    Jun 22, 2026Nils Grandien, David Steinmann, Felix Friedrich +1Hierarchical Representation LearningSparse Autoencoders

  21. Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders

    Jun 19, 2026Sergio Lanza, Jae Hee Lee, Stefan WermterVision-Language ModelsVLM Interpretability

  22. From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

    Jun 16, 2026Dibyanayan Bandyopadhyay, Asif EkbalLLM InterpretabilitySparse Autoencoders

  23. SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

    Jun 16, 2026Mingyue Cui, Linghui Shen, Xingyi YangSparse AutoencodersMechanistic Interpretability

  24. Analyzing Visual Aircraft Representations with Sparse Autoencoders

    Jun 13, 2026Deepshik SharmaVisual Representation LearningSparse Autoencoders

  25. Size Doesn't Matter: Cosine-Scored Sparse Autoencoders

    Jun 13, 2026Silen Naihin, Lev StamblerRepresentation LearningSparse Autoencoders