cs.LGSep 30, 2026

Whitening Improves Robustness to Spurious Correlations in Linear Probes

Authors: Floris Holstege, Bram Wouters, Noud van Giersbergen, Cees Diks

Organizations: University of Amsterdam, Department of Quantitative Economics · Tinbergen Institute

Abstract

Deep neural networks tend to rely on simple features that may be spurious and thus fail to generalize. We study this problem in the setting of linear probes, where a (generalized) linear model is fitted on the representations of a (pretrained) model. We use the connection of these models to the max-margin classifier, and show they favor directions associated with large eigenvalues of the covariance matrix. Whitening removes this preference by equalizing the eigenvalues of the covariance matrix. This observation motivates whitening as a preprocessing step that can reduce reliance on spurious correlations without requiring prior knowledge of their presence or labeled data. We examine the effect of whitening on a synthetic data-generating process and standard spurious correlation benchmarks, and find that it improves robustness. We also find that whitening can improve robustness when added to existing approaches.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Mar 2, 2026cs.LG

Spectral Overfitting in Noisy Linear Probing of Pretrained Representations

Frozen pretrained features are often treated as a safe interface for downstream learning: only a small linear readout is trained, while the backbone is fixed. We show that this readout can still overfit noisy labels in a structured way. A label-blind PCA rank sweep reveals a sharp spectral pattern: under label noise, exposing all pretrained directions can hurt clean accuracy, and intermediate ranks often recover much of the lost performance. Rank-matched random projections help less, and measured between-class signal is strongly concentrated in leading PCs. The pattern appears across three ImageNet-pretrained backbones on CIFAR-10, with gains up to 36.0±0.836.0\pm0.8 points over the default full-rank probe at 40% noise. Tuned full-rank probes outperform validation-selected PCA probes, so we present the sweep as a diagnostic of spectral overfitting rather than a competitive noisy-label method.
Jun 17, 2026cs.LG

Comparing Linear Probes with Mahalanobis Cosine Similarity

Linear probes are widely used in interpretability research and often compared by cosine similarity. The Mahalanobis cosine similarity (MCS) between two directions, which reweights the inner product by test data covariance, is a natural task-aware refinement. Ying et al. (2026) report that a probe's MCS to a reference probe trained on the out-of-distribution (OOD) data near-perfectly linearly predicts the probe's OOD AUROC (R^2 = 0.98). Here, we extend this empirical finding across models, layers, and concept domains, and prove this general phenomenon in closed form: For balanced classes whose projections are Gaussian, OOD AUROC and MCS to the reference probe are linear because both are sigmoid-shaped functions of the probe's signal-to-noise ratio (SNR) on the test data. The theory also predicts when this linearity fails, which we verify empirically. MCS offers a theoretically grounded and empirically effective alternative to Euclidean cosine similarity for comparing linear probes.
Jun 1, 2026cs.LG

Mitigating Spurious Correlations with Memorization-Guided Dataset De-Biasing

Real-world datasets often contain spurious correlations that are not causally related to the target label. When such correlations dominate the majority of training samples, models tend to rely on them, leading to misclassification of minority samples that do not exhibit the same spurious patterns. While a potential approach is to select subsets of data to better represent the minority samples, this may require access to group labels, which are typically unknown. Furthermore, as we demonstrate, widely used sample scoring functions in the invariant subset or coreset selection literature largely depend on spurious features and therefore fail to accurately capture the importance or difficulty of core, causally relevant features. Accordingly, we propose to mitigate spurious correlations by developing a two-stage sample scoring function that disentangles the learning dynamics of core and spurious features and evaluates their difficulty separately. Based on our proposed metric, we introduce a new algorithm to find and prioritize informative samples both with and without spurious correlations. Extensive experiments demonstrate that a standard ERM model trained on our selected samples achieves superior performance compared to state-of-the-art debiasing techniques, while requiring as little as 10% of the original training data.