When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models
Organizations: Department of Engineering Science University of Oxford
Abstract
Sparse autoencoders (SAEs) decompose dense model activations into discrete latents, making individual features easy to interpret--and easy to misinterpret. In EEG foundation models, this creates a tempting inference: if removing alpha-band activity strongly changes a latent's activation, one might conclude that the latent represents alpha activity. Across 27 settings spanning three backbones, three EEG datasets, and three network depths, this interpretation initially appears compelling: alpha removal changes latent firing 7.3 times more than an equal-width sham notch (95% CI [6.2, 8.7], bootstrapped over settings). However, the alpha filter also deletes far more signal than the sham. After normalizing by removed spectral energy, the ratio falls to 0.28 (95% CI [0.22, 0.36]) and exceeds one in none of the 27 settings. Latents selected for their response to alpha removal are, on clean EEG, slightly anti-correlated with relative alpha power (mean r = -0.073), giving no support for a simple alpha-detector reading. Motivated by this failure case, we propose a validation ladder for semantic interpretations of SAE latents: it asks in turn whether a latent responds, whether that response survives controlling for how much signal the intervention removes, whether it is specific rather than broadly fragile, and whether the proposed property is visible on unperturbed data--while separately testing the stronger claim that the latent matters to a task classifier. Perturbation sensitivity alone does not establish what an SAE latent represents.
Figures & tables
| Evidence | Claim supported |
|---|---|
| Responds to alpha removal | Alpha-sensitive latent |
| Response survives distortion control and competing perturbations | Alpha-selective latent |
| Above, and independently covaries with alpha structure in clean EEG | Evidence for an alpha-representing latent |
| Above, and readout intervention changes task behavior | Task-relevant alpha representation |
| ratio, 27 settings | Mean | 95% CI | Settings |
|---|---|---|---|
| raw | 7.33 | 27/27 | |
| per unit time-domain distortion | 1.32 | 17/27 | |
| per unit removed spectral energy | 0.28 | 0/27 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Jaccard | Ratio | ||||
|---|---|---|---|---|---|
| Model | sham | sham | /sham | ||
| CBraMod | .541 | .819 | .00338 | .00046 | 7.4 |
| REVE | .466 | .841 | .00127 | .00014 | 8.9 |
| DINOv3 | .484 | .782 | .00169 | .00033 | 5.1 |