Sparse Autoencoders

Also known as SAE

Momentum

16 papers in the last four weeks, up 100% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 149

All topics
CardsList
  1. Rational Sparse Autoencoder

    Jun 12, 2026Naiyu Yin, Yue YuActivation SparsityLLM Interpretability

  2. Decompose Sparsely Where You Should, Absorb Densely Where You Should No

    Jun 12, 2026Ruixuan Deng, Zehao Jin, Zekun Wang +1Transformer InterpretabilityRepresentation Learning

  3. Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

    Jun 8, 2026Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1TTS SynthesisLanguage Model Steering

  4. A Unifying Framework for Concept-Based Representational Similarity

    Jun 8, 2026Grégoire Dhimoïla, Victor Boutin, Agustin Martin Picard +2Representational Similarity AnalysisCross-Modal Alignment

  5. Interactions Between Crosscoder Features: A Compact Proofs Perspective

    Jun 8, 2026Dmitry Manning-Coe, Thomas Read, Anna Soligo +4Feature Interaction ModelingSparse Autoencoders

  6. Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

    Jun 5, 2026Georgii Aparin, Vadim Popov, Tasnima Sadekova +1Audio-Language Model Hallucination DetectionSparse Autoencoders

  7. A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders

    Jun 5, 2026Chenhao Zhang, Chris Lin, Su-In LeeSparse AutoencodersNeural Network Interpretability

  8. Interpreting Brain Responses to Language with Sparse Features from Language Models

    Jun 5, 2026Michael A. Lepori, Kendrick Kay, Greta TuckuteLLM InterpretabilitySparse Autoencoders

  9. Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision Transformers

    Jun 4, 2026Tang Li, Yanlin Chen, Mengmeng Ma +1Transformer InterpretabilityVision Transformer

  10. Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design

    Jun 1, 2026Michał Brzozowski, Neo Christopher ChungSparse AutoencodersMechanistic Interpretability

  11. Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

    May 31, 2026Hongfei Du, Jiacheng Shi, Sidi Lu +2TTS SynthesisSparse Autoencoders

  12. Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

    May 30, 2026Hwiyeong Lee, Ingyu Bang, Uiji Hwang +2Representation LearningLLM Interpretability

  13. On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders

    May 29, 2026Elana Simon, Etowah Adams, James ZouActivation OutliersSparse Autoencoders

  14. Toward Identifiable Sparse Autoencoders

    May 29, 2026Walter Nelson, Theofanis Karaletsos, Francesco LocatelloParameter IdentifiabilitySparse Autoencoders

  15. Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines

    May 29, 2026Mikkel Godsk Jørgensen, Lars Kai HansenLanguage Model SteeringLLM Interpretability

  16. Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing

    May 27, 2026Li Lei, Madalina Ciobanu, Qingqing Mao +1Task VectorsLLM Interpretability

  17. Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression

    May 27, 2026Tue M. Cao, Nguyen Do, My T. ThaiSparse AutoencodersMechanistic Interpretability

  18. ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions

    May 27, 2026Prathyush Poduval, Calvin Yeung, Neel Desai +1Transformer InterpretabilityActivation Sparsity

  19. Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models

    May 27, 2026Calvin Yeung, Prathyush Poduval, Ali Zakeri +2AutoencodersSparse Autoencoders