cs.CVSep 27, 2026

Position Aware Layer Queries for Test Time Training in Vision Language Models

Authors: Rajat Modi, Priyank Pathak, Xin Liang, Yogesh Singh Rawat

Organizations: Institute of Artificial Intelligence University of Central Florida Orlando, FL 32816, USA

Abstract

Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, (OOD)) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM, incurring substantial computational cost. We observe that one forward pass with all the intermediate layer outputs already yields far more signal than the final embedding from all augmentations. We introduce Layer Query Network (LQN), a lightweight approach that can adapt a frozen VLM (teacher) in a single forward pass of the VLM via a small model (student). LQN uses Position-Aware Distillation (PAD) to mimic the teacher VLM's intermediate-layer spatial tokens by querying spatial coordinates of intermediate tokens. LQN additionally relies on Location Consistency Regularization (LCR), a self-supervision technique, replacing expensive O(H x W) image augmentation with O(1) coordinate sampling. Integrating these, LQN i) adapts and improves zero-shot CLIP ViT-B/16 by 9.8% Top-1 on OOD ImageNet, ii) outperforms the previous best GS-Bias on fine-grained classification by 3.9% Top-1, iii) achieves faster convergence than TPS for CLIP ResNet-50 (47 mins vs 55 mins), iv) generalizes adaptation to VLMs like SigLIP, EVA-CLIP, and CoCa, and lightweight students like MLP, ResNet, VGG, and v) extends to panoptic, instance, and semantic segmentation.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Local Margin Restoration for Test-Time Adaptation of Vision-Language Models

    Aug 3, 2026Yan Huang, Guowei Wang, Xu Wang +2Vision-Language Model AdaptationTest-Time Adaptation

  2. Prototype-Based Test-Time Adaptation of Vision-Language Models

    Apr 23, 2026Zhaohong Huang, Yuxin Zhang, Wenjing Liu +2Vision-Language Model AdaptationTest-Time Adaptation

  3. Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models

    May 19, 2026Jia-Wei Hai, Yijun Wang, Xiu-Shen WeiVision-Language Model AdaptationTest-Time Adaptation