Sparse Autoencoders

Also known as SAE

Momentum

16 papers in the last four weeks, up 100% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 149

All topics
CardsList
  1. Universal Boosts, Specific Suppressors: Sparse Autoencoder Steering of Medical Vision-Language Models

    May 24, 2026Farhad Nooralahzadeh, Benjamin Gundersen, Nicolas Deperrois +7Radiology Report GenerationHealthcare

  2. Steered Generation via Gradient-Based Optimization on Sparse Query Features

    May 21, 2026Sumanta Bhattacharyya, Pedram RooshenasConstrained Motion PlanningLanguage Model Steering

  3. Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection

    May 21, 2026Yusser Al Ghussin, Daniil Gurgurov, Tanja Baeumel +3Multilingual Language ModelsLanguage Model Steering

  4. Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

    May 21, 2026Dongxin Guo, Jikun Wu, Siu Ming YiuNeural ProcessesLLM Interpretability

  5. Reading Task Failure Off the Activations: A Sparse-Feature Audit of GPT-2 Small on Indirect Object Identification

    May 21, 2026Mahdi NasermoghadasiTransformer InterpretabilitySparse Autoencoders

  6. Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models

    May 21, 2026Piotr Kubaty, Patryk Marszałek, Łukasz Struski +3Disentangled Representation LearningMultimodal Disentangled Representation Learning

  7. Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)

    May 18, 2026Michał Brzozowski, Neo Christopher ChungSparse AutoencodersMechanistic Interpretability

  8. Beyond Linear Superposition: Discovering Climate Features in AI Weather Models with KAN-SAE

    May 17, 2026Minjong CheonTransformer InterpretabilityClimate Science

  9. Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies

    May 17, 2026Xinchen Jin, Aditya Chatterjee, Pranav Kumar +1VLM InterpretabilitySparse Autoencoders

  10. Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models

    May 15, 2026Fabian Morelli, Arnas Uselis, Ankit Sonthalia +1Distribution Shift RobustnessVLM Robustness

  11. The Rate-Distortion-Polysemanticity Tradeoff in SAEs

    May 14, 2026Tommaso Mencattini, Francesco Montagna, Francesco LocatelloRate-Distortion TheoryRepresentation Disentanglement

  12. Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders

    May 13, 2026William Lehn-Schiøler, Magnus Ruud Kjær, Rahul Thapa +10Sparse AutoencodersMechanistic Interpretability

  13. Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning

    May 12, 2026Hyeonjin Kim, Hangyeol Jung, Heechan Yun +2Disentangled Representation LearningSparse Autoencoders

  14. HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds

    May 11, 2026Honghan Wu, Tianyan Wang, Jiacong Mi +2Hierarchical Representation LearningSparse Autoencoders

  15. The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws

    May 11, 2026Eslam Zaher, Maciej Trzaskowski, Quan Nguyen +1Neural Representation GeometrySparse Autoencoders

  16. fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery

    May 10, 2026Andreas D. Demou, Panagiotis Koromilas, James Oldfield +2Transformer InterpretabilitySparse Autoencoders

  17. SMIXAE: Towards Unsupervised Manifold Discovery in Language Models

    May 9, 2026Collin FrancelAutoencodersLLM Interpretability

  18. Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders

    May 8, 2026Tue M. Cao, Hoang X. Nhat, Raed Alharbi +2Hierarchical Representation LearningSparse Autoencoders

  19. Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

    May 8, 2026Hao Wang, Yiqun Sun, Pengfei Wei +2VLM RobustnessAdversarial Attacks on VLMs

  20. SoftSAE: Dynamic Top-K Selection for Adaptive Sparse Autoencoders

    May 7, 2026Jakub Stępień, Marcin Mazur, Jacek Tabor +1Sparse AutoencodersMechanistic Interpretability

  21. Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders

    May 7, 2026Shunchang Liu, Xin Chen, Belen Martin Urcelay +1Reward ModelingSparse Autoencoders

  22. From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features

    May 7, 2026Ruben Fernandez-Boullon, Pablo Magariños-Docampo, Javier Perez-RoblesTransformer InterpretabilitySparse Autoencoders