cs.CVOct 5, 2026

Efficient Test-time Adaptation through Candidate Verification and Divergence Shifts

Authors: Seungmin Oh, Seunghun Kang, Jongbin Ryu

Organizations: Ajou University

Abstract

Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes, caches, priors, or feature statistics, often incurring additional computational overhead. In this paper, we take a different perspective and reframe VLM-TTA as candidate verification rather than prediction adjustment. We propose Test-Time Correction (TTC), a hypothesis-based correction framework guided by a simple principle: hypothesize, reconstruct, correct. Given a test feature and its top-k candidate labels, TTC treats each candidate label as a hypothesis, reconstructs the feature within the corresponding latent subspace stored in a memory bank, and measures the resulting divergence shift. This shift quantifies how much the candidate subspace and its relations to other candidates change after the hypothetical insertion of the test feature. A correct candidate hypothesis induces only a small shift, whereas an incorrect one perturbs the subspace more strongly. TTC therefore corrects the prediction by selecting the candidate with the minimum aggregated divergence shift. This training-free candidate-verification mechanism avoids iterative optimization and provides a favorable accuracy-efficiency trade-off. Across five TTA settings and 15 benchmark datasets, including zero-shot classification, domain generalization, few-shot classification, base-to-novel generalization, and cross-dataset evaluation, TTC consistently improves accuracy over state-of-the-art VLM-TTA methods while achieving up to 2x speedup, over 3x lower CPU memory usage, and up to 1.4x lower GPU memory usage than the lowest-memory training-free baseline.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Prototype-Based Test-Time Adaptation of Vision-Language Models

    Apr 23, 2026Zhaohong Huang, Yuxin Zhang, Wenjing Liu +2Vision-Language Model AdaptationTest-Time Adaptation

  2. Local Margin Restoration for Test-Time Adaptation of Vision-Language Models

    Aug 3, 2026Yan Huang, Guowei Wang, Xu Wang +2Vision-Language Model AdaptationTest-Time Adaptation

  3. To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models

    Sep 8, 2026Siru Jiang, Yuwei Liang, Jian Liang +2Vision-Language Model AdaptationTest-Time Adaptation