cs.LGSep 28, 2026

When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models

Authors: Yucong Cao, Chenqi Li, Tingting Zhu

Organizations: Department of Engineering Science University of Oxford

Abstract

Sparse autoencoders (SAEs) decompose dense model activations into discrete latents, making individual features easy to interpret--and easy to misinterpret. In EEG foundation models, this creates a tempting inference: if removing alpha-band activity strongly changes a latent's activation, one might conclude that the latent represents alpha activity. Across 27 settings spanning three backbones, three EEG datasets, and three network depths, this interpretation initially appears compelling: alpha removal changes latent firing 7.3 times more than an equal-width sham notch (95% CI [6.2, 8.7], bootstrapped over settings). However, the alpha filter also deletes far more signal than the sham. After normalizing by removed spectral energy, the ratio falls to 0.28 (95% CI [0.22, 0.36]) and exceeds one in none of the 27 settings. Latents selected for their response to alpha removal are, on clean EEG, slightly anti-correlated with relative alpha power (mean r = -0.073), giving no support for a simple alpha-detector reading. Motivated by this failure case, we propose a validation ladder for semantic interpretations of SAE latents: it asks in turn whether a latent responds, whether that response survives controlling for how much signal the intervention removes, whether it is specific rather than broadly fragile, and whether the proposed property is visible on unperturbed data--while separately testing the stronger claim that the latent matters to a task classifier. Perturbation sensitivity alone does not establish what an SAE latent represents.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders

    May 13, 2026William Lehn-Schiøler, Magnus Ruud Kjær, Rahul Thapa +10Electroencephalography Foundation ModelsImproving Sparse Autoencoders

  2. What Do EEG Foundation Models Capture from Human Brain Signals?

    May 12, 2026Ling Tang, Qian Chen, Jilin Mei +6Electroencephalography Foundation ModelsElectroencephalography

  3. Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

    Jul 27, 2026Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty +2Sparse Autoencoder FeaturesInterpretability