Sparse Autoencoders

Also known as SAE

Momentum

16 papers in the last four weeks, up 100% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 149

All topics
CardsList
  1. Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders

    Oct 7, 2026Beimnet Bekele Guta, Xiaoyu Yang, Guangzhi Sun +1Disentangled Representation LearningSelf-Supervised Speech Representation Learning

  2. Inference and learning in sparse autoencoders as natural gradient flow

    Oct 5, 2026Hadi Vafaii, Tejas Rao, David Chanin +6Gradient DescentSparse Autoencoders

  3. Backdooring Sparse Autoencoders

    Oct 5, 2026Enrico Ahlers, Daniel Passon, Tobias Kiecker +2LLM Backdoor AttacksAdversarial Attacks

  4. A Testable Theory of Atomic Features

    Oct 5, 2026Kenny Peng, Jon Kleinberg, Nikhil GargSparse AutoencodersRepresentation Geometry in Language Models

  5. From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment

    Oct 1, 2026Liwei Lin, Gus XiaRepresentation LearningSparse Autoencoders

  6. D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders

    Sep 30, 2026Xinyue Xu, Jiahao Zhang, Lijie Hu +2Transformer InterpretabilityDiffusion Transformer

  7. Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

    Sep 29, 2026Zhenting Huang, Bo Jiang, Junnan Liu +2Activation SparsityLLM Interpretability

  8. Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery

    Sep 28, 2026Xu Wang, Yifan Yang, TingHao YU +1Representation LearningLLM Interpretability

  9. From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features

    Sep 28, 2026Dewen Liu, Zixuan Li, Jonathan Pan +4Sparse AutoencodersMechanistic Interpretability

  10. Verifying the Linear Representation Hypothesis: How Interpretable Are Vision SAEs?

    Sep 28, 2026Teodor Chiaburu, Franz Motzkus, Frank Haußer +1Linear Representation HypothesisSparse Autoencoders

  11. When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models

    Sep 28, 2026Yucong Cao, Chenqi Li, Tingting ZhuPerturbation-Based Feature AttributionSparse Autoencoders

  12. Parts-of-Speech as Emergent Categories in SAE Latent Space

    Sep 24, 2026Alessandro Bondielli, Lucia Passaro, Serena Auriemma +1Sparse AutoencodersRepresentation Geometry in Language Models

  13. Local Sparsity Enables Unsupervised LLM Safety Detection

    Sep 17, 2026Xin Chen, Gil Kur, Alexander Shevchenko +1Sparse AutoencodersLLM Safety

  14. What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

    Sep 14, 2026Lingheng Du, Yiming Tang, Xufeng Duan +1RL for Language ModelsSparse Autoencoders

  15. Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery

    Sep 14, 2026Hendrik Droste, Christian Medeiros Adriano, Kathrin Korte +1Transformer InterpretabilitySparse Autoencoders

  16. Tracing Stereotypes from Representation to Output in Multilingual LLMs

    Sep 8, 2026Ariun-Erdene Tumurchuluun, Yusser Al Ghussin, Pinzhen Chen +2Multilingual Language ModelsSocial Bias in Language Models

  17. EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders

    Sep 3, 2026Xinghao Wang, Dong Li, Wei Yu +5Video Diffusion ModelsSparse Autoencoders

  18. Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment

    Sep 1, 2026Mian Zhong, Katherine A. Keith, Anjalie FieldCausal Effect EstimationSparse Autoencoders

  19. SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

    Aug 13, 2026Weihan Meng, Hongzhu Guo, Yi Jing +5LLM InterpretabilitySparse Autoencoders

  20. Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

    Aug 11, 2026Nikolai Bolik, Lennart Stöpler, Artur AndrzejakRepresentation LearningLLM Interpretability

  21. Measuring Semantic Abstractness of SAE Features via Nonlocality

    Aug 11, 2026Chuqiao Lin, Shivaji Sondhi, Xiao-Liang QiSparse AutoencodersMechanistic Interpretability

  22. Multimodal Model Diffing for Feature Discovery and Control

    Aug 10, 2026Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4Sparse AutoencodersMultimodal Model Interpretability

  23. Steering dense music retrieval with open-vocabulary concept discovery

    Aug 9, 2026Julien Guinot, Alain Riou, Elio Quinton +1Feature AttributionSparse Autoencoders

  24. Interpretable GOHR Agents via Sparse Autoencoders

    Jul 27, 2026Shiwei Tan, Yusong Zhao, Weiyi Qin +6Transformer InterpretabilitySparse Autoencoders