Technical note on: Zero-Training Feature-Space Alignment via Information Geometry
Authors: Behraj Khan, Tahir Qasim Syed, Syed Ahmad Chan Bukhari
Organizations: School of Mathematics and Computer Science, Institute of Business Administration Karachi, Pakistan · Division of Computer Science, Mathematics and Science, St. John’s University, USA
Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness under covariate shift without modifying model parameters. ZFGA is based on the observation that distribution shifts distort feature-space geometry. It estimates the Fisher information matrix of the predictive distribution with respect to feature embeddings and applies a linear transformation that aligns test-feature Fisher geometry with a reference geometry computed from clean data. This provides a natural-gradient-inspired preconditioning step in feature space. We evaluate ZFGA on CIFAR-10-C and ImageNet-C using ResNet-50, DINO ViT-S/16, and CLIP ViT-B/32. ZFGA consistently improves over zero-shot inference across all three models, although it is not the strongest method for every model. Covariance whitening performs better on ResNet-50, while Fisher whitening is statistically indistinguishable from ZFGA on CLIP. Across six training-free and gradient-based alternatives (covariance whitening, Fisher whitening, TENT, T3A, LAME, and AdaNPC), ZFGA is the only method that does not substantially harm any of the three model families. The Fisher geometry distortion is also positively correlated with ZFGA gain (Pearson r = 0.366, p = 0.017), providing preliminary evidence that geometric misalignment contributes to robustness degradation. ZFGA requires only forward passes and matrix operations at inference time, offering a lightweight and deterministic alternative to optimization-based test-time adaptation.
Figures & tables
Figure 1: Overview of Zero-Training Fisher Geometry Alignment (ZFGA). Clean reference images and corrupted test images are processed by a frozen encoder z=fθ(x) to obtain feature embeddings. The corresponding predictive distributions yield empirical Fisher information matrices I^ref and I^te , whose Frobenius discrepancy ΔF=∥I^ref−I^te∥F quantifies geometric distortion under covariate shift. ZFGA constructs the closed-form alignment matrix A=(I^te+ϵI)−1/2(I^ref+ϵI)1/2 and transforms the test embedding as z′=Az . Classification then uses the fixed temperature-scaled cosine model p(y∣x)∝exp(τ,z′⊤ty) with frozen prototypes ty . Zero-training: frozen encoder and prototypes, no parameter updates, and only forward passes and matrix operations at inference time.
Model
Frozen
Cov.
Fisher
ZFGA
TENT / TENT-LN
ResNet-50
Acc. ± std
68.50 ± 0.30
49.54 ± 1.17
65.49 ± 0.97
74.08 ± 0.26
73.31 ± 0.70
Δ
–
−18.96
−3.01
+5.58
+4.81
DINOv3 ViT-S/16
Acc. ± std
84.89 ± 0.74
33.76 ± 1.21
88.20 ± 0.92
84.97 ± 0.69
84.89 ± 0.74
Δ
–
−51.13
+3.31
+0.08
+0.00
Table 1: Main Results on CIFAR-10-C (Severities 2–3). Mean accuracy (%) ± std over 3 seeds, averaged over 7 corruption types. Δ is gain over the frozen baseline. Best non-TENT result per model in bold . DINO/CLIP use TENT-LN (LayerNorm variant); see Appendix .3 .
Model
Zero-shot
T3A
LAME
AdaNPC
ZFGA
ResNet-50
59.53 ± 0.75
60.16 ± 0.60
58.08 ± 0.60
64.49 ± 0.76
60.25 ± 0.74
DINOv3 ViT-S/16
67.60 ± 1.10
66.18 ± 0.99
64.02 ± 0.97
69.66 ± 0.78
67.74 ± 1.06
CLIP ViT-B/32
78.70 ± 0.98
79.57 ± 0.48
77.93 ± 1.32
74.52 ± 1.38
78.98 ± 1.02
Table 2: Comparison with training-free/backprop-free baselines on CIFAR-10-C (Severities 2–3). Mean accuracy (%) ± std over 3 seeds. All methods evaluated on identical encoders, corruption subset, and severities as Table 1 .
Model
Zero-shot
Cov. Whit.
Fisher Whit.
ZFGA
ResNet-50
27.82
29.80
28.88
30.18
DINOv3 ViT-S/16
38.18
15.61(−59.1%)
37.77
38.44
CLIP ViT-B/32
54.45
32.49(−40.3%)
54.82
55.02
Table 3: Results on ImageNet-C (Severity 3 only, restricted to 3 corruption types). Mean top-1 accuracy (%) averaged over Gaussian Noise, Motion Blur, and Contrast; single-run point estimates, no seed variance reported (see Section 4.2 ). Best results in bold .
Method
ResNet-50
DINOv3 ViT-S/16
CLIP ViT-B/32
Frozen baseline
750 ± 135 ms
6315 ± 122 ms
1590 ± 41 ms
Cov. Whitening
1537 ± 459 ms
6356 ± 129 ms
1610 ± 43 ms
Fisher Whitening
1645 ± 708 ms
6322 ± 121 ms
—
ZFGA
1565 ± 370 ms
6324 ± 120 ms
1606 ± 42 ms
TENT / TENT-LN
17405 ± 839 ms
19545 ± 250 ms
3717 ± 42 ms
TENT adapted params
53,120 (0.226%)
19,200 (0.089%)
39,936 (0.045%)
Table 4: Inference latency per 512-sample batch (mean ± std over 5 runs, single GPU). Adapted-parameter counts are for the gradient-based baseline.
Figure 2: Correlation between Fisher distortion and ZFGA effectiveness. Across models and corruptions (severities 2–3), larger geometric distortion (x-axis: ΔF ) is weakly associated with greater ZFGA gain over zero-shot baseline (y-axis). Pearson r=0.366 , p=0.017 ( r2≈0.13 ); this pooled correlation is preliminary evidence and should not be read as a strong or precisely quantified effect.
Figure 3: Fisher matrix eigenvalue spectrum. Top 50 eigenvalues for clean (green solid) vs. corrupted (red dashed) Fisher matrices on Gaussian noise severity 2. ResNet shows the largest spectrum shift (highest condition number) of the three models shown, DINO an intermediate shift, and CLIP the smallest shift. This ordering is descriptive of the three models studied and is not independently statistically tested.
Figure 4: Ablation studies. (Left) Batch size sensitivity: ZFGA accuracy on CLIP with Gaussian noise (sev. 2) stabilizes around n=128 –256. (Middle) Regularization ϵ : stable across [10−6,10−3] , showing robustness to hyperparameter choice. (Right) Severity sensitivity: ZFGA gain vs. corruption severity for Gaussian noise only (not averaged over all seven CIFAR-10-C corruption types, and not directly comparable to the multi-corruption headline numbers in Table 1 ). ResNet shows an inverted-U pattern peaking at severity 3 (+4.30%), where geometric distortion appears maximal without catastrophic information loss. Comparatively robust models (DINO, CLIP) show flatter profiles.
Corruption
Sev.
Zero-shot
Cov.
Fisher
TENT
T3A
LAME
AdaNPC
ZFGA
Gaussian noise
2
42.2
48.2
38.2
43.0
43.4
40.8
54.3
43.4
Gaussian noise
3
31.2
35.0
28.3
31.2
33.1
27.1
41.6
32.0
Motion blur
2
57.8
59.6
58.5
57.1
58.3
56.1
56.5
58.3
Motion blur
3
46.5
48.5
48.9
46.5
49.8
45.7
45.2
47.9
Defocus blur
2
70.8
74.1
69.4
71.2
69.1
67.6
76.5
71.1
Defocus blur
3
61.1
65.9
63.0
60.6
63.0
61.3
64.7
61.8
Table 5: ResNet-50: per-corruption accuracy (%) on CIFAR-10-C, severities 2–3. Mean over 3 seeds per cell.
Corruption
Sev.
Zero-shot
Cov.
Fisher
TENT
T3A
LAME
AdaNPC
ZFGA
Gaussian noise
2
41.3
21.7
47.7
40.5
49.9
30.5
23.7
42.0
Gaussian noise
3
26.0
16.7
31.4
26.0
39.4
17.5
11.3
26.8
Motion blur
2
68.0
13.0
65.1
69.3
64.1
65.9
73.6
68.1
Motion blur
3
63.0
12.3
59.1
64.5
59.4
59.4
65.2
63.1
Defocus blur
2
76.9
16.9
74.7
76.4
72.3
73.1
80.5
77.0
Defocus blur
3
71.4
16.1
68.8
72.5
67.0
68.2
73.8
71.5
Table 6: DINOv3 ViT-S/16: per-corruption accuracy (%) on CIFAR-10-C, severities 2–3. Mean over 3 seeds per cell.
Corruption
Sev.
Zero-shot
Cov.
Fisher
TENT
T3A
LAME
AdaNPC
ZFGA
Gaussian noise
2
59.5
26.0
62.7
60.2
60.0
52.1
22.1
61.7
Gaussian noise
3
45.8
26.8
50.8
46.6
46.7
34.2
12.0
48.2
Motion blur
2
79.0
39.6
80.6
79.1
80.7
80.1
81.4
79.0
Motion blur
3
71.3
34.2
73.0
72.1
73.5
65.2
70.0
71.7
Defocus blur
2
86.3
38.4
86.2
87.0
88.7
88.7
89.8
86.2
Defocus blur
3
83.5
39.3
83.7
84.1
85.4
85.3
86.6
83.7
Table 7: CLIP ViT-B/32: per-corruption accuracy (%) on CIFAR-10-C, severities 2–3. Mean over 3 seeds per cell.
1NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences · School of Computer Science and Engineering, Central South University · University of Chinese Academy of Sciences