Sparse Autoencoders

Also known as SAE

Momentum

16 papers in the last four weeks, up 100% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 149

All topics
CardsList
  1. Feature Starvation as Geometric Instability in Sparse Autoencoders

    May 6, 2026Faris Chaudhry, Keisuke Yano, Anthea MonodRepresentation LearningSparse Autoencoders

  2. Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes

    May 4, 2026Michael A. Riegler, Birk Sebastian Frostelid Torpmann-HagenSparse AutoencodersMechanistic Interpretability

  3. GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models

    May 3, 2026Favour Nerrise, Lucy Yin, Mohammad H. Abbasi +2Alzheimer's DiseaseMedical Imaging Foundation Models

  4. Do Sparse Autoencoders Capture Concept Manifolds?

    Apr 30, 2026Usha Bhalla, Thomas Fel, Can Rager +9Neural Representation GeometrySparse Autoencoders

  5. MoRFI: Monotonic Sparse Autoencoder Feature Identification

    Apr 29, 2026Dimitris Dimakopoulos, Shay B. Cohen, Ioannis KonstasLLM Hallucination MitigationLLM Interpretability

  6. Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection

    Apr 29, 2026Ahyoung Oh, Wonseok Shin, Songkuk KimSparse AutoencodersOOD Detection

  7. A Unifying Framework for Unsupervised Concept Extraction

    Apr 27, 2026Chandler Squires, Pradeep RavikumarUnsupervised LearningParameter Identifiability

  8. Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

    Apr 26, 2026John Winnicki, Abeynaya Gnanasekaran, Eric DarveKG ConstructionLLM Interpretability

  9. From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks

    Apr 25, 2026Nilanjana Das, Mathew Dawit, Aman Chadha +1LLM InterpretabilitySparse Autoencoders

  10. From Tokens to Concepts: Leveraging SAE for SPLADE

    Apr 23, 2026Yuxuan Zong, Mathias Vast, Basile Van Cooten +2Learned Sparse RetrievalSparse Autoencoders

  11. Towards Understanding the Robustness of Sparse Autoencoders

    Apr 20, 2026Ahson Saiyed, Sabrina Sadiekh, Chirag AgarwalAdversarial Attacks on LLMsJailbreak Robustness

  12. Structural Instability of Feature Composition

    Apr 18, 2026Yunpeng ZhouFeature SuperpositionTransformer Interpretability

  13. Improving Sparse Autoencoder with Dynamic Attention

    Apr 16, 2026Dongsheng Wang, Jinsen Zhang, Dawei Su +1Sparse AutoencodersNeural Network Interpretability

  14. Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation

    Mar 30, 2026Vitória Barin-Pacela, Shruti Joshi, Isabela Camacho +2Sparse AutoencodersSparse Coding

  15. Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information

    Mar 18, 2026Shih-Heng Wang, Tiantian Feng, Aditya Kommineni +4Neural Audio CodecsSparse Autoencoders

  16. Step-Level Sparse Autoencoder for Reasoning Process Interpretation

    Mar 3, 2026Xuan Yang, Jiayu Liu, Yuhang Lai +3LLM InterpretabilitySparse Autoencoders

  17. SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data

    Feb 16, 2026David Chanin, Adrià Garriga-AlonsoSparse AutoencodersSynthetic Benchmark Generation

  18. Sparse Autoencoders are Capable LLM Jailbreak Mitigators

    Feb 12, 2026Yannick Assogba, Jacopo Cortellazzi, Javier Abad +3Sparse AutoencodersLLM Jailbreak Attacks

  19. Variational Sparse Paired Autoencoders (vsPAIR) for Inverse Problems and Uncertainty Quantification

    Feb 3, 2026Jack Michael Solomon, Rishi Leburu, Matthias ChungVariational AutoencodersUncertainty Quantification

  20. Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Reasoning

    Feb 2, 2026Yadong Wang, Haodong Chen, Yu Tian +3Sparse AutoencodersContinuous Latent Reasoning

  21. Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders

    Nov 21, 2025Jacob Beattie, Samuel Stevens, Neil Rosser +2Vision Foundation ModelsSparse Autoencoders

  22. Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

    Aug 22, 2025David Chanin, Adrià Garriga-AlonsoActivation SparsityLLM Interpretability

  23. Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages

    Jul 15, 2025Lyzander Marciano Andrylie, Inaya Rahmanisa, Mahardika Krisna Ihsani +3Multilingual Language ModelsRepresentation Learning

  24. Position: Use Sparse Autoencoders to Discover Unknowns

    Jun 30, 2025Kenny Peng, Rajiv Movva, Jon Kleinberg +2Explainable Artificial IntelligenceSparse Autoencoders

  25. SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization

    Date pendingJingyi He, Haiyan Zhao, Ruxue Shi +4LLM InterpretabilitySparse Autoencoders

  26. SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

    Date pendingYuqiao Tan, Shizhu He, Jun Zhao +1AI Agent EvaluationSparse Autoencoders