Technical note on: Zero-Training Feature-Space Alignment via Information Geometry
Authors: Behraj Khan, Tahir Qasim Syed, Syed Ahmad Chan Bukhari
Organizations: School of Mathematics and Computer Science, Institute of Business Administration Karachi, Pakistan · Division of Computer Science, Mathematics and Science, St. John’s University, USA
Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness under covariate shift without modifying model parameters. ZFGA is based on the observation that distribution shifts distort feature-space geometry. It estimates the Fisher information matrix of the predictive distribution with respect to feature embeddings and applies a linear transformation that aligns test-feature Fisher geometry with a reference geometry computed from clean data. This provides a natural-gradient-inspired preconditioning step in feature space. We evaluate ZFGA on CIFAR-10-C and ImageNet-C using ResNet-50, DINO ViT-S/16, and CLIP ViT-B/32. ZFGA consistently improves over zero-shot inference across all three models, although it is not the strongest method for every model. Covariance whitening performs better on ResNet-50, while Fisher whitening is statistically indistinguishable from ZFGA on CLIP. Across six training-free and gradient-based alternatives (covariance whitening, Fisher whitening, TENT, T3A, LAME, and AdaNPC), ZFGA is the only method that does not substantially harm any of the three model families. The Fisher geometry distortion is also positively correlated with ZFGA gain (Pearson r = 0.366, p = 0.017), providing preliminary evidence that geometric misalignment contributes to robustness degradation. ZFGA requires only forward passes and matrix operations at inference time, offering a lightweight and deterministic alternative to optimization-based test-time adaptation.
Figures & tables
Figure 1: Overview of Zero-Training Fisher Geometry Alignment (ZFGA). Clean reference images and corrupted test images are processed by a frozen encoder z=fθ(x) to obtain feature embeddings. The corresponding predictive distributions yield empirical Fisher information matrices I^ref and I^te , whose Frobenius discrepancy ΔF=∥I^ref−I^te∥F quantifies geometric distortion under covariate shift. ZFGA constructs the closed-form alignment matrix A=(I^te+ϵI)−1/2(I^ref+ϵI)1/2 and transforms the test embedding as z′=Az . Classification then uses the fixed temperature-scaled cosine model p(y∣x)∝exp(τ,z′⊤ty) with frozen prototypes ty . Zero-training: frozen encoder and prototypes, no parameter updates, and only forward passes and matrix operations at inference time.
Model
Frozen
Cov.
Fisher
ZFGA
TENT / TENT-LN
ResNet-50
Acc. ± std
68.50 ± 0.30
49.54 ± 1.17
65.49 ± 0.97
74.08 ± 0.26
73.31 ± 0.70
Δ
–
−18.96
−3.01
+5.58
+4.81
DINOv3 ViT-S/16
Acc. ± std
84.89 ± 0.74
33.76 ± 1.21
88.20 ± 0.92
84.97 ± 0.69
84.89 ± 0.74
Δ
–
−51.13
+3.31
+0.08
+0.00
Table 1: Main Results on CIFAR-10-C (Severities 2–3). Mean accuracy (%) ± std over 3 seeds, averaged over 7 corruption types. Δ is gain over the frozen baseline. Best non-TENT result per model in bold . DINO/CLIP use TENT-LN (LayerNorm variant); see Appendix .3 .
Model
Zero-shot
T3A
LAME
AdaNPC
ZFGA
ResNet-50
59.53 ± 0.75
60.16 ± 0.60
58.08 ± 0.60
64.49 ± 0.76
60.25 ± 0.74
DINOv3 ViT-S/16
67.60 ± 1.10
66.18 ± 0.99
64.02 ± 0.97
69.66 ± 0.78
67.74 ± 1.06
CLIP ViT-B/32
78.70 ± 0.98
79.57 ± 0.48
77.93 ± 1.32
74.52 ± 1.38
78.98 ± 1.02
Table 2: Comparison with training-free/backprop-free baselines on CIFAR-10-C (Severities 2–3). Mean accuracy (%) ± std over 3 seeds. All methods evaluated on identical encoders, corruption subset, and severities as Table 1 .
Model
Zero-shot
Cov. Whit.
Fisher Whit.
ZFGA
ResNet-50
27.82
29.80
28.88
30.18
DINOv3 ViT-S/16
38.18
15.61(−59.1%)
37.77
38.44
CLIP ViT-B/32
54.45
32.49(−40.3%)
54.82
55.02
Table 3: Results on ImageNet-C (Severity 3 only, restricted to 3 corruption types). Mean top-1 accuracy (%) averaged over Gaussian Noise, Motion Blur, and Contrast; single-run point estimates, no seed variance reported (see Section 4.2 ). Best results in bold .
Method
ResNet-50
DINOv3 ViT-S/16
CLIP ViT-B/32
Frozen baseline
750 ± 135 ms
6315 ± 122 ms
1590 ± 41 ms
Cov. Whitening
1537 ± 459 ms
6356 ± 129 ms
1610 ± 43 ms
Fisher Whitening
1645 ± 708 ms
6322 ± 121 ms
—
ZFGA
1565 ± 370 ms
6324 ± 120 ms
1606 ± 42 ms
TENT / TENT-LN
17405 ± 839 ms
19545 ± 250 ms
3717 ± 42 ms
TENT adapted params
53,120 (0.226%)
19,200 (0.089%)
39,936 (0.045%)
Table 4: Inference latency per 512-sample batch (mean ± std over 5 runs, single GPU). Adapted-parameter counts are for the gradient-based baseline.
Figure 2: Correlation between Fisher distortion and ZFGA effectiveness. Across models and corruptions (severities 2–3), larger geometric distortion (x-axis: ΔF ) is weakly associated with greater ZFGA gain over zero-shot baseline (y-axis). Pearson r=0.366 , p=0.017 ( r2≈0.13 ); this pooled correlation is preliminary evidence and should not be read as a strong or precisely quantified effect.
Figure 3: Fisher matrix eigenvalue spectrum. Top 50 eigenvalues for clean (green solid) vs. corrupted (red dashed) Fisher matrices on Gaussian noise severity 2. ResNet shows the largest spectrum shift (highest condition number) of the three models shown, DINO an intermediate shift, and CLIP the smallest shift. This ordering is descriptive of the three models studied and is not independently statistically tested.
Figure 4: Ablation studies. (Left) Batch size sensitivity: ZFGA accuracy on CLIP with Gaussian noise (sev. 2) stabilizes around n=128 –256. (Middle) Regularization ϵ : stable across [10−6,10−3] , showing robustness to hyperparameter choice. (Right) Severity sensitivity: ZFGA gain vs. corruption severity for Gaussian noise only (not averaged over all seven CIFAR-10-C corruption types, and not directly comparable to the multi-corruption headline numbers in Table 1 ). ResNet shows an inverted-U pattern peaking at severity 3 (+4.30%), where geometric distortion appears maximal without catastrophic information loss. Comparatively robust models (DINO, CLIP) show flatter profiles.
Corruption
Sev.
Zero-shot
Cov.
Fisher
TENT
T3A
LAME
AdaNPC
ZFGA
Gaussian noise
2
42.2
48.2
38.2
43.0
43.4
40.8
54.3
43.4
Gaussian noise
3
31.2
35.0
28.3
31.2
33.1
27.1
41.6
32.0
Motion blur
2
57.8
59.6
58.5
57.1
58.3
56.1
56.5
58.3
Motion blur
3
46.5
48.5
48.9
46.5
49.8
45.7
45.2
47.9
Defocus blur
2
70.8
74.1
69.4
71.2
69.1
67.6
76.5
71.1
Defocus blur
3
61.1
65.9
63.0
60.6
63.0
61.3
64.7
61.8
Table 5: ResNet-50: per-corruption accuracy (%) on CIFAR-10-C, severities 2–3. Mean over 3 seeds per cell.
Corruption
Sev.
Zero-shot
Cov.
Fisher
TENT
T3A
LAME
AdaNPC
ZFGA
Gaussian noise
2
41.3
21.7
47.7
40.5
49.9
30.5
23.7
42.0
Gaussian noise
3
26.0
16.7
31.4
26.0
39.4
17.5
11.3
26.8
Motion blur
2
68.0
13.0
65.1
69.3
64.1
65.9
73.6
68.1
Motion blur
3
63.0
12.3
59.1
64.5
59.4
59.4
65.2
63.1
Defocus blur
2
76.9
16.9
74.7
76.4
72.3
73.1
80.5
77.0
Defocus blur
3
71.4
16.1
68.8
72.5
67.0
68.2
73.8
71.5
Table 6: DINOv3 ViT-S/16: per-corruption accuracy (%) on CIFAR-10-C, severities 2–3. Mean over 3 seeds per cell.
Corruption
Sev.
Zero-shot
Cov.
Fisher
TENT
T3A
LAME
AdaNPC
ZFGA
Gaussian noise
2
59.5
26.0
62.7
60.2
60.0
52.1
22.1
61.7
Gaussian noise
3
45.8
26.8
50.8
46.6
46.7
34.2
12.0
48.2
Motion blur
2
79.0
39.6
80.6
79.1
80.7
80.1
81.4
79.0
Motion blur
3
71.3
34.2
73.0
72.1
73.5
65.2
70.0
71.7
Defocus blur
2
86.3
38.4
86.2
87.0
88.7
88.7
89.8
86.2
Defocus blur
3
83.5
39.3
83.7
84.1
85.4
85.3
86.6
83.7
Table 7: CLIP ViT-B/32: per-corruption accuracy (%) on CIFAR-10-C, severities 2–3. Mean over 3 seeds per cell.
While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution (OOD) outliers prevalent in real-world applications. This discrepancy motivates Noisy TTA (NTTA), an online task to filter noisy OOD samples on the fly while maximizing in-distribution (ID) classification accuracy. Existing zero-shot NTTA approaches typically rely on test-time discriminative training, leading to overconfident misclassifications and significantly degraded inference efficiency. To address these limitations, we propose a novel framework named Dual Distribution Estimation (DDE), shifting the zero-shot NTTA paradigm from instance-level learning to training-free Gaussian distribution modeling. DDE incorporates two novel modules: Positive Feature Distribution Estimation (PFDE) and Negative Label Distribution Estimation (NLDE). PFDE explicitly models class-wise inclusion and exclusion Gaussian distributions to formulate a calibrated contrastive score, robustly enhancing ID accuracy. In parallel, NLDE improves OOD identification by explicitly modeling the negative label distribution to mine highly discriminative labels, effectively mitigating spurious correlations. Extensive experiments show that on the large-scale ImageNet benchmark, DDE achieves an improvement of 3.70% in harmonic mean accuracy and reduces the FPR95 for OOD detection by 6.20%, while ensuring highly scalable and efficient online inference. Furthermore, DDE is zero-shot and training-free, demonstrating remarkable robustness in data-scarce scenarios. Codes are available at https://github.com/ZhuWenjie98/DDE.
Wenjie Zhu, Yabin Zhang, Liang Xu +3
The Hong Kong Polytechnic University · Harbin Institute of Technology (Shenzhen) · Shanghai Jiao Tong University +1
Pretrained vision-language models such as CLIP exhibit strong zero-shot generalization but remain sensitive to distribution shifts. Test-time adaptation adapts models during inference without access to source data or target labels, offering a practical way to handle such shifts. However, existing methods typically assume that test samples come from a single, consistent domain, while in practice, test data often include samples from mixed domains with distinct characteristics. Consequently, their performance degrades under mixed-domain settings. To address this, we present Ramen, a framework for robust test-time adaptation through active sample selection. For each incoming test sample, Ramen retrieves a customized batch of relevant samples from previously seen data based on two criteria: domain consistency, which ensures that adaptation focuses on data from similar domains, and prediction balance, which mitigates adaptation bias caused by skewed predictions. To improve efficiency, Ramen employs an embedding-gradient cache that stores the embeddings and sample-level gradients of past test images. The stored embeddings are used to retrieve relevant samples, and the corresponding gradients are aggregated for model updates, eliminating the need for any additional forward or backward passes. Our theoretical analysis provides insight into why the proposed adaptation mechanism is effective under mixed-domain shifts. Experiments on multiple image corruption and domain-shift benchmarks demonstrate that Ramen achieves strong and consistent performance, offering robust and efficient adaptation in complex mixed-domain scenarios. Our code is available at https://github.com/baowenxuan/Ramen .
Vision-language models like CLIP have demonstrated remarkable zero-shot transfer capabilities. However, their susceptibility to imperceptible adversarial perturbations remains a critical security concern. While test-time defenses offer a pragmatic solution for deployed models, existing approaches typically rely on gradient-based optimization during inference, incurring significant computational overhead. In this paper, we revisit the role of data augmentation in CLIP robustness and observe that augmentations are not equally effective: specific augmentations consistently provide robust geometric cues that align with correct class semantics in the hyperspherical feature space. Based on this, we propose Adaptive Geodesic Correction (AGC), a training-free defense mechanism that requires no parameter updates. AGC identifies a reliable augmentation as a geometric anchor and corrects the input feature towards it, utilizing an adaptive step size to balance robustness against clean accuracy preservation. AGC achieves superior performance across eight fine-grained datasets and three CLIP backbones, improving average robust accuracy by 44.4% over state-of-the-art baseline while delivering a 10× reduction in inference latency. Our findings reveal a fundamental geometric property of CLIP features, offering a highly efficient and effective paradigm for robust multimodal deployment.
Zhiwei Li, Jiacheng Xue, Weining Wang +4
1NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences · School of Computer Science and Engineering, Central South University · University of Chinese Academy of Sciences