cs.LGSep 28, 2026

Interference Beyond Geometry in Concept Extraction

Authors: Valérie Costa, Bahareh Tolooshams

Organizations: University of Alberta, Amii · Harvard University · University of Alberta, Amii Canada CIFAR AI Chair

Abstract

Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions, distinguishing constructive from destructive interference and frequent weak interactions from rare strong ones. Under local fixed-support assumptions, we characterize how architectural constraints shape interference through four mechanisms: feature orthogonalization, bias compensation, gain adaptation, and encoder-decoder separation. Experiments with sparse autoencoders show that constrained architectures selectively reduce overlap among co-active features, while bias, gain, and encoder freedom allow constructive cross-contributions to remain. Together, these results show that interference in learned representations depends not only on feature geometry, but also on how features are used and on the architecture that produces their codes.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Structural Instability of Feature Composition

    Apr 18, 2026Yunpeng ZhouFeature LearningSteering

  2. High-probability guarantees for linear accessibility in feature superposition

    Sep 10, 2026Enrico VompaConcentration BoundsInterference

  3. Towards Isolated Interventions via Almost Orthogonal Features in Language Models

    Feb 4, 2026Moritz Miller, Florent Draye, Bernhard SchölkopfLatent Feature InterventionsMechanistic Interpretability