Improving Sparse Autoencoders

Latest papers 47

All topics
CardsList
  1. Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

    Sep 29, 2026Zhenting Huang, Bo Jiang, Junnan Liu +2Improving Sparse AutoencodersTop-K

  2. Identifying ODEs from Unstructured Data with Causal Representation Learning

    Sep 29, 2026Alessandro Trenta, Riccardo Massidda, Davide Bacciu +1Ordinary Differential EquationsCausal Representation Learning

  3. Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery

    Sep 28, 2026Xu Wang, Yifan Yang, TingHao YU +1Improving Sparse AutoencodersDeep Semantic Representations

  4. Verifying the Linear Representation Hypothesis: How Interpretable Are Vision SAEs?

    Sep 28, 2026Teodor Chiaburu, Franz Motzkus, Frank Haußer +1Improving Sparse AutoencodersInterpretability

  5. Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

    Sep 16, 2026Devesh Tiwari, Camille Davis, Shivank Sinha +3Model ActivationsCausal Reasoning

  6. What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

    Sep 14, 2026Lingheng Du, Yiming Tang, Xufeng Duan +1Large Language Model Reinforcement LearningImproving Sparse Autoencoders

  7. MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

    Aug 31, 2026Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang +5InterpretabilityMechanistic Interpretability

  8. PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders

    Aug 11, 2026Man Jiang, Ouxiang Li, Weibao Xue +4Concept ErasureText-To-Image Diffusion Models

  9. Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

    Jul 28, 2026Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal +1Model ActivationsLarge Language Model Safety

  10. Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?

    Jul 19, 2026Aniket DeshpandeImproving Sparse AutoencodersReconstruction Error

  11. Persistent Sparse Autoencoders: Learning Feature-Specific Timescales in Language Model Representations

    Jul 19, 2026Haoyan Luo, Mateo Espinosa Zarlenga, Mateja JamnikImproving Sparse AutoencodersFuture Latent Representations

  12. When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities

    Jul 9, 2026Weiduo Liao, Yunqiao Yang, Ying WeiImproving Sparse AutoencodersModalities

  13. Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

    Jul 9, 2026Bendegúz Váradi, Zoltán KmettyImproving Sparse AutoencodersBert-Based Models

  14. Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

    Jun 26, 2026Kevin Der, Harish Kamath, Ben ThompsonSparse Autoencoder FeaturesImproving Sparse Autoencoders

  15. VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

    Jun 26, 2026Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah +1Sparse Autoencoder FeaturesImproving Sparse Autoencoders

  16. Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

    Jun 25, 2026Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. HosseiniImproving Sparse AutoencodersSparsity

  17. Steering Vision-Language Models with Joint Sparse Autoencoders

    Jun 24, 2026Huizhen Shu, Xuying Li, Hongxu Lin +2Improving Sparse AutoencodersCross-Modal Attention

  18. Effects of sparsity and superposition on loss in simple autoencoders

    Jun 16, 2026Mriganka Basu Roy Chowdhury, Eric McLaughlin WeinerImproving Sparse AutoencodersAutoencoder Architectures

  19. Rational Sparse Autoencoder

    Jun 12, 2026Naiyu Yin, Yue YuImproving Sparse AutoencodersAutoencoder Architectures

  20. Decompose Sparsely Where You Should, Absorb Densely Where You Should No

    Jun 12, 2026Ruixuan Deng, Zehao Jin, Zekun Wang +1Improving Sparse AutoencodersSparsity

  21. Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design

    Jun 1, 2026Michał Brzozowski, Neo Christopher ChungImproving Sparse AutoencodersUnsupervised Dictionary Learning

  22. Toward Identifiable Sparse Autoencoders

    May 29, 2026Walter Nelson, Theofanis Karaletsos, Francesco LocatelloImproving Sparse AutoencodersSparse Autoencoder Features

  23. ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions

    May 27, 2026Prathyush Poduval, Calvin Yeung, Neel Desai +1Improving Sparse AutoencodersTransformer Residual Streams

  24. Universal Boosts, Specific Suppressors: Sparse Autoencoder Steering of Medical Vision-Language Models

    May 24, 2026Farhad Nooralahzadeh, Benjamin Gundersen, Nicolas Deperrois +7Medical Vision-Language ModelsImproving Sparse Autoencoders

  25. Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies

    May 17, 2026Xinchen Jin, Aditya Chatterjee, Pranav Kumar +1InterpretabilityImproving Sparse Autoencoders

  26. Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders

    May 13, 2026William Lehn-Schiøler, Magnus Ruud Kjær, Rahul Thapa +10Electroencephalography Foundation ModelsImproving Sparse Autoencoders

  27. Domain Restriction via Multi SAE Layer Transitions

    May 12, 2026Elias Shaheen, Avi MendelsonDomain-Specific LanguageImproving Sparse Autoencoders

  28. Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect

    May 11, 2026Joao Paulo Cavalcante Presa, Savio Salvarino Teles de OliveiraEmotionSelf

  29. The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws

    May 11, 2026Eslam Zaher, Maciej Trzaskowski, Quan Nguyen +1Improving Sparse AutoencodersScaling Laws

  30. Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders

    May 8, 2026Tue M. Cao, Hoang X. Nhat, Raed Alharbi +2Improving Sparse AutoencodersHierarchical

  31. Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes

    May 4, 2026Michael A. Riegler, Birk Sebastian Frostelid Torpmann-HagenImproving Sparse AutoencodersModel Activations

  32. GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models

    May 3, 2026Favour Nerrise, Lucy Yin, Mohammad H. Abbasi +2Improving Sparse AutoencodersBiomarker

  33. MoRFI: Monotonic Sparse Autoencoder Feature Identification

    Apr 29, 2026Dimitris Dimakopoulos, Shay B. Cohen, Ioannis KonstasLarge Language Model HallucinationModel Fine-Tuning

  34. Improving Sparse Autoencoder with Dynamic Attention

    Apr 16, 2026Dongsheng Wang, Jinsen Zhang, Dawei Su +1Improving Sparse AutoencodersDynamic Sparse Attention

  35. Step-Level Sparse Autoencoder for Reasoning Process Interpretation

    Mar 3, 2026Xuan Yang, Jiayu Liu, Yuhang Lai +3LLM Reasoning StrategiesImproving Sparse Autoencoders

  36. SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data

    Feb 16, 2026David Chanin, Adrià Garriga-AlonsoSparse Autoencoder FeaturesImproving Sparse Autoencoders

  37. Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

    Aug 22, 2025David Chanin, Adrià Garriga-AlonsoImproving Sparse AutoencodersModel Activations

  38. Position: Use Sparse Autoencoders to Discover Unknowns

    Jun 30, 2025Kenny Peng, Rajiv Movva, Jon Kleinberg +2Improving Sparse AutoencodersModel Discovery