Sparse Autoencoder Features

Latest papers 45

All topics
CardsList
  1. A Testable Theory of Atomic Features

    Oct 5, 2026Kenny Peng, Jon Kleinberg, Nikhil GargSparse Autoencoder Features

  2. When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models

    Sep 28, 2026Yucong Cao, Chenqi Li, Tingting ZhuElectroencephalography Foundation ModelsSparse Autoencoder Features

  3. Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery

    Sep 14, 2026Hendrik Droste, Christian Medeiros Adriano, Kathrin Korte +1Sparse Autoencoder FeaturesFeature Learning

  4. Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering

    Sep 8, 2026Wenbo Zhang, Zhongxiang Sun, Zhiguang Han +1Sparse Autoencoder FeaturesSteering

  5. SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

    Aug 13, 2026Weihan Meng, Hongzhu Guo, Yi Jing +5Sparse Autoencoder FeaturesVerbalization

  6. Measuring Semantic Abstractness of SAE Features via Nonlocality

    Aug 11, 2026Chuqiao Lin, Shivaji Sondhi, Xiao-Liang QiSparse Autoencoder FeaturesMechanistic Interpretability

  7. CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

    Aug 6, 2026Mehrshad Saadatinia, Parsa Razmara, Ardalan Aryashad +2Linear Activation SteeringSteering

  8. Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

    Jul 27, 2026Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty +2Sparse Autoencoder FeaturesInterpretability

  9. Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

    Jul 22, 2026Seonglae Cho, Zekun Wu, Kleyton Da Costa +3Sparse Autoencoder FeaturesModel Activations

  10. Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

    Jun 26, 2026Kevin Der, Harish Kamath, Ben ThompsonSparse Autoencoder FeaturesImproving Sparse Autoencoders

  11. VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

    Jun 26, 2026Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah +1Sparse Autoencoder FeaturesImproving Sparse Autoencoders

  12. At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

    Jun 24, 2026Praneet Suresh, Jack Stanley, Sonia Joseph +2Out-Of-DistributionTransformer Architectures

  13. VFUSE: Virulent Feature Understanding with Sparse autoEncoders

    Jun 8, 2026Michael Yu, Matthew L. OlsonProtein DesignSparse Autoencoder Features

  14. Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

    Jun 5, 2026Chaitanya Shinde, Hadi Hajieghrary, Paul Schmitt +3Automated VehiclesControllability

  15. On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders

    May 29, 2026Elana Simon, Etowah Adams, James ZouSparse Autoencoder FeaturesModel Activations

  16. Toward Identifiable Sparse Autoencoders

    May 29, 2026Walter Nelson, Theofanis Karaletsos, Francesco LocatelloImproving Sparse AutoencodersSparse Autoencoder Features

  17. Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression

    May 27, 2026Tue M. Cao, Nguyen Do, My T. ThaiSparse Autoencoder FeaturesDistance

  18. Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

    May 21, 2026Dongxin Guo, Jikun Wu, Siu Ming YiuSparse Autoencoder FeaturesHuman Visual Cortex

  19. Reading Task Failure Off the Activations: A Sparse-Feature Audit of GPT-2 Small on Indirect Object Identification

    May 21, 2026Mahdi NasermoghadasiSparse Autoencoder FeaturesSpurious Correlations

  20. Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)

    May 18, 2026Michał Brzozowski, Neo Christopher ChungSparse Autoencoder FeaturesFeature Learning

  21. On the Interpretability of Whisper Encodings Using Sparse Autoencoders

    May 12, 2026Dan Pluth, Zachary Nicholas Houghton, Yu Zhou +1Speech EncoderSparse Autoencoder Features

  22. SMIXAE: Towards Unsupervised Manifold Discovery in Language Models

    May 9, 2026Collin FrancelSparse Autoencoder FeaturesModel Activations

  23. From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features

    May 7, 2026Ruben Fernandez-Boullon, Pablo Magariños-Docampo, Javier Perez-RoblesSparse Autoencoder FeaturesToken Co-Occurrence Graphs

  24. Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

    Apr 26, 2026John Winnicki, Abeynaya Gnanasekaran, Eric DarveSparse Autoencoder FeaturesKnowledge Graphs

  25. From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks

    Apr 25, 2026Nilanjana Das, Mathew Dawit, Aman Chadha +1Jailbreak Success RatesUnsafe

  26. SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data

    Feb 16, 2026David Chanin, Adrià Garriga-AlonsoSparse Autoencoder FeaturesImproving Sparse Autoencoders

  27. Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages

    Jul 15, 2025Lyzander Marciano Andrylie, Inaya Rahmanisa, Mahardika Krisna Ihsani +3Sparse Autoencoder FeaturesMultilingual Language Models

  28. SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization

    Date pendingJingyi He, Haiyan Zhao, Ruxue Shi +4Sparse Autoencoder FeaturesExplainable AI Methods