Position Aware Layer Queries for Test Time Training in Vision Language Models
Organizations: Institute of Artificial Intelligence University of Central Florida Orlando, FL 32816, USA
Abstract
Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, (OOD)) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM, incurring substantial computational cost. We observe that one forward pass with all the intermediate layer outputs already yields far more signal than the final embedding from all augmentations. We introduce Layer Query Network (LQN), a lightweight approach that can adapt a frozen VLM (teacher) in a single forward pass of the VLM via a small model (student). LQN uses Position-Aware Distillation (PAD) to mimic the teacher VLM's intermediate-layer spatial tokens by querying spatial coordinates of intermediate tokens. LQN additionally relies on Location Consistency Regularization (LCR), a self-supervision technique, replacing expensive O(H x W) image augmentation with O(1) coordinate sampling. Integrating these, LQN i) adapts and improves zero-shot CLIP ViT-B/16 by 9.8% Top-1 on OOD ImageNet, ii) outperforms the previous best GS-Bias on fine-grained classification by 3.9% Top-1, iii) achieves faster convergence than TPS for CLIP ResNet-50 (47 mins vs 55 mins), iv) generalizes adaptation to VLMs like SigLIP, EVA-CLIP, and CoCa, and lightweight students like MLP, ResNet, VGG, and v) extends to panoptic, instance, and semantic segmentation.
Figures & tables
| Method | Augment. | ImageNet | ImageNet-A | ImageNet-V2 | ImageNet-R | ImageNet-Sketch | Avg | OOD Avg | |
| CLIP-B/16 | CLIP-ViT-B/16 | ✗ | 66.7 | 47.8 | 60.8 | 73.9 | 46.0 | 59.0 | 57.1 |
| Ensemble | ✗ | 68.3 | 49.8 | 61.8 | 77.6 | 48.2 | 61.1 | 59.4 | |
| TPT | ✓ | 68.9 | 54.7 | 63.4 | 77.0 | 47.9 | 62.4 | 60.8 | |
| Diff-TPT | ✓ | 70.3 | 55.6 | 65.1 | 75.0 | 46.8 | 62.6 | 60.6 | |
| MTA + TPT | ✓ | 70.0 | 58.0 | 64.2 | 78.3 | 49.6 | 64.0 | 62.5 | |
| APM | ✗ | 68.1 | 52.1 | 67.2 | 76.5 | 49.3 | 62.6 | 61.3 | |
| Method | Augment. | Flower102 | DTD | Pets | UCF101 | Caltech101 | Food101 | SUN397 | Aircraft | EuroSAT | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|
| CLIP-ViT-B/16 | ✗ | 67.4 | 44.3 | 88.3 | 65.1 | 93.4 | 83.7 | 62.6 | 23.7 | 42.0 | 63.4 |
| Ensemble | ✗ | 67.0 | 45.0 | 86.9 | 65.2 | 93.6 | 82.9 | 65.6 | 23.2 | 50.4 | 64.4 |
| TPT | ✓ | 69.0 | 47.8 | 87.8 | 68.0 | 94.2 | 84.7 | 65.5 | 24.8 | 42.4 | 64.9 |
| DiffTPT | ✓ | 70.1 | 47.0 | 88.2 | 62.6 | 92.4 | 87.2 | 65.7 | 25.6 | 43.1 | 64.7 |
| MTA | ✓ | 68.0 | 45.9 | 88.2 | 68.6 | 94.2 | 85.0 | 66.6 | 25.2 | 45.3 | 65.2 |
| APM | ✗ | 62.0 | 48.9 | 81.6 | 72.6 | 89.6 | 84.2 | 65.7 | 29.7 | 55.7 | 65.5 |
| Panoptic | Instance | Semantic | |||
| COCO | ADE20K | COCO | Cityscapes | ADE20K | |
| Method | PQ ↑ | PQ ↑ | AP ↑ | mIoU ↑ | mIoU ↑ |
| EoMT | 58.3 | 51.7 | 48.8 | 84.2 | 58.4 |
| LQN (Our) | 64.5 | 57.1 | 55.2 | 87.5 | 64.3 |
| LQN (Pretrained, Our ) | 64.9 | 58.6 | 56.7 | 88.2 | 65.8 |
| ImageNet and OOD | Fine-grained classification | ||||||||||||||||
| Method | ImageNet | IN-A | IN-V2 | IN-R | IN-Sketch | Avg | OOD Avg | Flowers102 | DTD | Pets | UCF101 | Caltech101 | Food101 | SUN397 | Aircraft | EuroSAT | Avg |
| SigLIP | 76.0 | 45.3 | 68.9 | 90.3 | 67.9 | 69.7 | 68.1 | 85.8 | 64.7 | 94.1 | 72.5 | 90.5 | 89.8 | 69.8 | 43.8 | 43.8 | 72.8 |
| + LQN | 79.2 | 48.7 | 72.4 | 93.4 | 71.6 | 73.1 | 71.5 | 89.2 | 68.5 | 97.4 | 76.2 | 94.1 | 93.1 | 73.5 | 47.3 | 47.6 | 76.3 |
| EVA-CLIP | 76.1 | 64.6 | 68.9 | 89.1 | 63.3 | 72.4 | 71.5 | 72.0 | 59.3 | 93.7 | 74.1 | 90.4 | 89.7 | 71.9 | 28.5 | 69.9 | 72.2 |
| + LQN | 79.8 | 68.9 | 73.1 | 92.5 | 67.4 | 76.3 | 75.5 | 76.4 | 63.2 | 97.2 | 78.6 | 94.3 | 93.7 | 75.9 | 32.7 | 74.1 | 76.2 |
| CoCa | 63.6 | 21.5 | 55.7 | 73.2 | 51.3 | 53.1 | 50.4 | 64.7 | 53.3 | 89.1 | 61.4 | 89.1 | 77.3 | 66.1 | 18.8 | 45.3 | 62.8 |
| Parts | Food-101 | COCO |
|---|---|---|
| PAD (- ) | 71.8 | 54.9 |
| PAD + | 76.0 | 58.1 |
| PAD + LCR | 89.2 | 64.5 |
| Parts | Food-101 | COCO |
|---|---|---|
| PAD (- ) | 71.8 | 54.9 |
| PAD + | 76.0 | 58.1 |
| PAD + LCR | 89.2 | 64.5 |
| Student | Food-101 |
|---|---|
| VGG | 84.9 |
| ResNet 18 | 85.4 |
| ResNet 34 | 87.2 |
| MLP | 90.8 |
| Loss | Food-101 | COCO |
|---|---|---|
| L1 | 75.8 | 52.0 |
| MSE | 85.7 | 61.1 |
| Cosine | 89.2 | 64.5 |
| Pos. | Food-101 | COCO |
|---|---|---|
| Coord. | 33.1 | 27.9 |
| Sin-Cos | 88.6 | 63.8 |
| RoPE | 89.2 | 64.5 |
| Sampling Strategy | Food-101 | COCO |
|---|---|---|
| Sampling more from shallow layers | 81.9 | 43.7 |
| Sampling more from deeper layers | 87.8 | 52.0 |
| Uniform Sampling | 89.2 | 64.5 |
| Sampling Strategy | Food-101 | COCO |
|---|---|---|
| Sampling more from shallow layers | 81.9 | 43.7 |
| Sampling more from deeper layers | 87.8 | 52.0 |
| Uniform Sampling | 89.2 | 64.5 |
| Configuration | Food-101 | COCO |
|---|---|---|
| PAD + LCR (w/o weight Reg.) | 87.4 | 62.6 |
| PAD + LCR (w/ weight Reg.) | 89.2 | 64.5 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Requirements | ImageNet |
|---|---|---|
| TDA [CVPR’24] | History | 69.5 |
| DMN-ZS [CVPR’24] | History | 72.2 |
| DPE [NeurIPS’24] | History | 71.9 |
| DynaPrompt [ICLR’25] | History | 72.9 |
| CoOp [IJCV’22] | Labeled Data | 71.5 |
| CoCoOp [CVPR’22] | Labeled Data | 71.0 |
| Layer | Feature Dimension | Stride | Padding | ||
|---|---|---|---|---|---|
| (H W C) | Input / Output | ||||
| Input | |||||
| Encoder | Conv | 1 | 0 / 0 | ||
| Decoder | Linear | * | - | - | - |
| Linear | * | - | - | - | |
| Linear | * | - | - | - |
| Number of Test samples | 50000 (Imagenet Splits), variable for other datasets. |
|---|---|
| Testing iterations | 15 |
| Batch Size | 1 |
| Learning Rate | 1e-4 |
| Optimizer | Adam |
| Feature Output size | |
| Positional Encoding size |
| Dataset | ImageNet | Flowers102 | DTD | Caltech101 | Aircraft | IN-A | IN-v2 | Pets | Sun397 | Eurosat | IN-R | UCF101 | Food101 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.7 | 0.5 | 0.7 | 0.5 | 0.7 | 0.6 | 0.6 | 0.6 | 0.6 | 0.6 | 0.8 | 0.8 | 0.8 | |
| (in LCR) | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 | 1e-7 |
| N | 5 | 8 | 12 | 15 | 20 | 25 |
|---|---|---|---|---|---|---|
| Accuracy | 80.5 | 83.9 | 88.6 | 89.2 | 87.8 | 81.0 |
| Memory | ORIN | AGX Thor | Ampere | GFlops | |
|---|---|---|---|---|---|
| EoMT (teacher) | 2.1 GB | 0.5 sec | 0.1sec | 0.08sec | 4146 |
| LQN (student) / iter | 600 MB | 0.68sec | 0.37sec | 0.29sec | 24.5 |
| LQN total (15 iters) | 2.7 GB | 8.9sec | 4.3sec | 2.88sec | 4384 |
| Model | CLIP ViT-L | 2 Layers | 4 Layers | 8 Layers | 10 Layers |
|---|---|---|---|---|---|
| Accuracy | 74.9 | 75.8 | 76.5 | 77.3 | 77.0 |
| Model | Base | + LQN |
|---|---|---|
| Vision-Mamba-T | 78.3 | 80.9 |
| Vision-Mamba-S | 81.4 | 83.2 |
| P | brigh | cont | defoc | elast | fog | frost | gauss | glass | impul | jpeg | motn | pixel | shot | snow | zoom | Average | |
| Joint Train | ✓ | 62.3 | 4.5 | 26.7 | 39.9 | 25.7 | 30.0 | 5.8 | 16.3 | 5.8 | 45.3 | 30.9 | 45.9 | 7.1 | 25.1 | 31.8 | 24.8 |
| Fine-Tune | ✓ | 67.5 | 7.8 | 33.9 | 32.4 | 36.4 | 38.2 | 22.0 | 15.7 | 23.9 | 51.2 | 37.4 | 51.9 | 23.7 | 37.6 | 37.1 | 33.7 |
| ViT Probe | ✓ | 68.3 | 6.4 | 24.2 | 31.6 | 38.6 | 38.4 | 17.4 | 18.4 | 18.2 | 51.2 | 32.2 | 49.7 | 18.2 | 35.9 | 32.2 | 29.2 |
| TTT-MAE | ✓ | 69.1 | 9.8 | 34.4 | 50.7 | 44.7 | 50.7 | 30.5 | 36.9 | 32.4 | 63.0 | 41.9 | 63.0 | 33.0 | 42.8 | 45.9 | 44.4 |
| OpenCLIP VIT-L/14(t) | ✗ | 71.9 | 47.0 | 50.3 | 32.7 | 58.3 | 46.9 | 26.0 | 26.5 | 28.1 | 62.7 | 37.7 | 58.3 | 28.2 | 50.4 | 37.9 | 42.1 |
| APM | ✗ | 77.4 | 51.9 | 56.6 | 37.9 | 64.8 | 53.2 | 28.7 | 31.4 | 33.0 | 68.4 | 44.1 | 64.5 | 33.1 | 56.9 | 43.9 | 50.3 |
| P | brigh | cont | defoc | elast | fog | frost | gauss | glass | impul | jpeg | motn | pixel | shot | snow | zoom | Average | |
| Baseline | ✓ | 73.1 | 33.1 | 35.8 | 56.9 | 54.2 | 45.2 | 39.6 | 26.0 | 38.2 | 62.0 | 43.2 | 60.3 | 32.2 | 44.2 | 40.7 | 47.4 |
| TTT-MAE | ✓ | 72.7 | 39.6 | 45.7 | 64.9 | 58.3 | 52.6 | 48.5 | 42.8 | 47.6 | 67.0 | 50.5 | 66.6 | 42.4 | 45.7 | 51.5 | 53.2 |
| OpenCLIP VIT-L/14 | ✗ | 74.2 | 64.2 | 58.7 | 57.8 | 66.3 | 52.8 | 45.3 | 34.6 | 45.2 | 68.9 | 46.6 | 63.9 | 41.1 | 56.2 | 45.6 | 54.8 |
| APM | ✗ | 79.2 | 70.4 | 64.9 | 63.7 | 72.3 | 58.6 | 51.2 | 40.4 | 51.3 | 74.1 | 53.0 | 70.0 | 46.7 | 62.5 | 51.8 | 59.6 |
| LQN (Ours) | ✗ | 80.9 | 72.1 | 66.5 | 65.2 | 73.9 | 59.9 | 52.9 | 41.9 | 52.8 | 75.3 | 54.2 | 71.4 | 47.9 | 63.9 | 53.3 | 61.2 |
| P | brigh | cont | defoc | elast | fog | frost | gauss | glass | impul | jpeg | motn | pixel | shot | snow | zoom | Average | |
| Baseline | ✓ | 75.8 | 62.7 | 49.5 | 67.1 | 59.8 | 47.6 | 57.1 | 35.0 | 57.4 | 68.6 | 60.2 | 70.1 | 54.3 | 54.7 | 48.0 | 57.6 |
| TTT-MAE | ✓ | 75.8 | 64.4 | 59.4 | 71.2 | 64.0 | 54.0 | 63.6 | 50.7 | 64.2 | 71.3 | 64.2 | 73.1 | 61.8 | 58.0 | 57.4 | 64.4 |
| OpenCLIP VIT-L/14 | ✗ | 75.8 | 71.8 | 65.5 | 67.7 | 69.0 | 54.7 | 58.9 | 42.4 | 59.5 | 72.8 | 59.9 | 69.7 | 58.2 | 63.5 | 51.8 | 62.5 |
| APM | ✗ | 80.5 | 77.2 | 71.3 | 73.3 | 74.8 | 60.6 | 64.7 | 48.5 | 65.4 | 77.8 | 61.6 | 75.2 | 64.1 | 69.3 | 58.0 | 68.5 |
| LQN (Ours) | ✗ | 81.9 | 78.8 | 72.8 | 74.9 | 76.2 | 61.9 | 66.3 | 49.9 | 66.9 | 79.1 | 62.8 | 76.7 | 65.5 | 70.5 | 59.5 | 69.9 |
| P | brigh | cont | defoc | elast | fog | frost | gauss | glass | impul | jpeg | motn | pixel | shot | snow | zoom | Average | |
| Baseline | ✓ | 77.4 | 71.2 | 62.3 | 51.0 | 66.3 | 58.4 | 68.6 | 59.2 | 64.9 | 70.4 | 70.6 | 74.7 | 66.2 | 54.2 | 55.2 | 64.1 |
| TTT-MAE | ✓ | 77.8 | 71.5 | 69.4 | 49.7 | 69.8 | 62.7 | 72.5 | 66.4 | 70.0 | 72.7 | 72.3 | 76.2 | 70.6 | 58.7 | 63.6 | 68.3 |
| OpenCLIP VIT-L/14 | ✗ | 76.6 | 74.4 | 71.4 | 53.8 | 72.0 | 62.6 | 67.6 | 64.0 | 64.6 | 73.8 | 69.0 | 72.8 | 66.4 | 61.8 | 58.3 | 66.1 |
| APM | ✗ | 81.1 | 79.4 | 76.6 | 59.4 | 77.3 | 68.2 | 73.1 | 70.0 | 70.3 | 78.6 | 74.5 | 77.8 | 72.0 | 67.8 | 64.3 | 72.4 |
| LQN (Ours) | ✗ | 82.4 | 81.0 | 78.4 | 60.9 | 78.9 | 69.8 | 74.7 | 71.3 | 71.7 | 80.1 | 75.8 | 79.2 | 73.4 | 69.1 | 65.8 | 74.1 |
| P | brigh | cont | defoc | elast | fog | frost | gauss | glass | impul | jpeg | motn | pixel | shot | snow | zoom | Average | |
| Baseline | ✓ | 78.5 | 74.5 | 68.1 | 73.9 | 70.5 | 70.6 | 74.8 | 68.6 | 72.3 | 73.0 | 75.2 | 75.9 | 73.6 | 69.3 | 63.7 | 71.4 |
| TTT-MAE | ✓ | 78.9 | 74.7 | 72.5 | 74.7 | 72.9 | 72.2 | 76.8 | 72.2 | 75.5 | 74.5 | 75.8 | 77.0 | 75.9 | 71.9 | 69.3 | 73.1 |
| OpenCLIP VIT-L/14 | ✗ | 77.3 | 75.4 | 73.5 | 73.1 | 73.5 | 71.4 | 71.9 | 70.2 | 69.9 | 75.1 | 73.7 | 74.2 | 71.9 | 71.2 | 65.2 | 71.1 |
| APM | ✗ | 81.6 | 80.3 | 78.6 | 78.0 | 78.6 | 76.6 | 77.2 | 75.7 | 75.1 | 79.6 | 78.7 | 79.1 | 76.9 | 76.4 | 70.7 | 76.0 |
| LQN (Ours) | ✗ | 83.2 | 82.0 | 80.3 | 79.6 | 80.1 | 78.0 | 78.6 | 77.0 | 76.5 | 80.9 | 80.2 | 80.5 | 78.4 | 77.9 | 72.3 | 77.6 |
| Method | orig | gauss | shot | impul | defoc | glass | motn | zoom | snow | frost | fog | brit | contr | elas | pixel | jpeg | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TTT-Online | 8.2 | 25.8 | 22.6 | 30.6 | 14.6 | 34.4 | 18.3 | 17.1 | 20.0 | 18.0 | 16.9 | 11.2 | 15.6 | 21.6 | 18.1 | 21.2 | 19.1 |
| UDA-SS | 9.0 | 28.2 | 26.5 | 20.8 | 15.6 | 43.7 | 24.5 | 23.8 | 25.0 | 24.9 | 17.2 | 12.7 | 11.6 | 22.1 | 20.3 | 22.6 | 21.4 |
| Zeroshot | |||||||||||||||||
| CLIP ViT-L/14 | 4.63 | 35.4 | 32.3 | 21.9 | 19.3 | 49.7 | 19.3 | 17.3 | 17.0 | 15.1 | 21.6 | 8.4 | 15.9 | 34.6 | 25.0 | 27.4 | 24.5 |
| CLIP ViT-L/14 (t) | |||||||||||||||||
| APM | 3.5 | 21.9 | 30.1 | 13.7 | 15.2 | 34.1 | 11.9 | 11.1 | 15.0 | 9.0 | 13.5 | 5.8 | 9.5 | 23.0 | 15.8 | 17.0 | 14.8 |