Readout Blindness: VLM Scores Miss the Spatial Direction Their Frozen Encoders Retain
Organizations: ELLIS Institute Finland · Aalto University · University of Oulu · Nanyang Technological University
Abstract
CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when scoring rules treat the subject and object symmetrically, direction cancels regardless of encoder training. Guided by this analysis, we introduce Antisymmetric Displacement Readout (ADR), which aligns caption words with image patches in the frozen features and scores each relation by the signed displacement between matched object centroids. Notably, ADR succeeds without additional training or learned parameters, thereby demonstrating that directional information remains in the frozen encoder. However, text and world priors can inflate accuracy, so we further introduce prior deflation, which measures the benefit of the image-text pairing as the grounded gain over a null that pairs each item with an unrelated image. Extensive experiments across encoder families show that ADR substantially improves over deployed scores, which remain near chance on most direction-balanced sets even for fine-tuned encoders. Compared with more complex readouts, ADR outperforms the evaluated MLLM likelihood readouts and is competitive with their chat inference at a small fraction of the computation. These results support our claim that directional information can be recovered from frozen features by an appropriate readout. Our implementation and evaluation kit will be publicly available.
Figures & tables
| Subset | Random | Metric | CLIP | NegCLIP | SigLIP2 | LLaVA-1.5 | Qwen3.6 |
|---|---|---|---|---|---|---|---|
| likelihood | likelihood | ||||||
| Prior-based subset | 50.03 | acc. | 72.71 | 85.14 | 62.07 | 93.26 | 93.99 |
| VALSE actant swap | n.acc. | 41.31 | 62.80 | 28.35 | 88.51 | 88.41 | |
| dev. | |||||||
| gain | |||||||
| Direction-balanced subset | 50.01 | acc. | 49.64 | 50.26 | 49.95 | 65.80 | 74.76 |
| Readout | VG-Rel L/R | COCO | WU-A | WU-B | VSR | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | |
| Random | 50.01 | 50.01 | 0.0 | 49.99 | 49.99 | 0.0 | 25.02 | 25.02 | 0.0 | 50.01 | 50.01 | 0.0 | 50.01 | 50.01 | 0.0 |
| CLIP | 49.64 | 49.69 | 45.30 | 47.28 | 26.96 | 25.98 | 53.67 | 52.83 | 53.87 | 53.21 | |||||
| ADR- | 90.45 | 49.03 | 92.33 | 51.73 | 31.62 | 25.74 | 93.22 | 52.26 | 73.86 | 47.79 | |||||
| ADR- | 90.22 | 49.44 | 94.06 | 52.72 | 41.18 | 25.25 | 98.02 | 50.00 | 74.94 | 47.88 | |||||
| LAION-2B | 49.13 | 49.77 | 43.07 | 47.28 | 27.70 | 26.23 | 50.00 | 50.28 | 51.96 | 52.96 | |||||
| Readout | VG-Rel L/R | COCO | WU-A | WU-B | VSR | ||||||||||
| acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | |
| Deployed readouts | |||||||||||||||
| CLIP | 49.64 | 49.69 | 45.30 | 47.28 | 26.96 | 25.98 | 53.67 | 52.83 | 53.87 | 53.21 | |||||
| SigLIP2 | 49.95 | 50.54 | 49.01 | 45.05 | 29.17 | 24.75 | 51.69 | 48.87 | 54.70 | 52.96 | |||||
| Half-crop | |||||||||||||||
| CLIP | 83.67 | 49.16 | 82.92 | 55.45 | 44.61 | 22.30 | 100.00 | 50.00 | 70.52 | 47.88 | |||||
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| CLIP ViT-B/16 | SigLIP2-So400m | |
| Median text difference | 0.2251 | 0.2494 |
| Pairs with identical text features | 0 | 0 |
| Global score ties | 0 | 0 |
| Global real / null accuracy | 49.64 / 49.69 | 49.95 / 50.54 |
| Global wrong, ADR- correct, item count | 1785 | 1799 |
| Median text difference, global wrong and ADR- correct | 0.2264 | 0.2505 |
| Subset | random | metric | CLIP | NegCLIP | SigLIP2 | LLaVA-1.5 | Qwen3.6 |
|---|---|---|---|---|---|---|---|
| likelihood | likelihood | ||||||
| Direction-balanced subsets | |||||||
| VG-Rel left/right | 50.01 | acc. | 49.64 | 50.26 | 49.95 | 65.80 | 74.76 |
| n.acc. | 49.69 | 49.67 | 50.54 | 49.44 | 49.29 | ||
| dev. | |||||||
| gain | |||||||
| Readout | VG-Rel L/R | COCO | WU-A | WU-B | VSR | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | |
| random | 50.01 | 50.01 | 0.0 | 49.99 | 49.99 | 0.0 | 25.02 | 25.02 | 0.0 | 50.01 | 50.01 | 0.0 | 50.01 | 50.01 | 0.0 |
| frozen dual encoders, deployed score and ADR | |||||||||||||||
| CLIP ViT-B/16 | 49.64 | 49.69 | 45.30 | 47.28 | 26.96 | 25.98 | 53.67 | 52.83 | 53.87 | 53.21 | |||||
| ADR- | 90.45 | 49.03 | 92.33 | 51.73 | 31.62 | 25.74 | 93.22 | 52.26 | 73.86 | 47.79 | |||||
| readout | VG-Rel L/R | COCO | WU-B | VSR | pooled |
|---|---|---|---|---|---|
| (a) AUROC (derangement null) | |||||
| ADR- | 0.891 (0.510) | 0.904 (0.516) | 0.888 (0.449) | 0.713 (0.484) | 0.863 (0.504) |
| ADR- | 0.880 (0.509) | 0.899 (0.483) | 0.970 (0.506) | 0.694 (0.478) | 0.844 (0.504) |
| deployed cosine | 0.509 (0.505) | 0.461 (0.461) | 0.512 (0.481) | 0.504 (0.487) | 0.514 (0.504) |
| MLLM likelihood, Qwen3.6-27B | 0.717 (0.496) | 0.726 (0.555) | 0.829 (0.508) | 0.568 (0.461) | 0.661 (0.498) |
| subject noun | items | ViT-B/16 | SigLIP2 |
|---|---|---|---|
| bottle | 28 | 8 | 2 |
| bowl | 12 | 6 | 3 |
| cup | 12 | 6 | 0 |
| orange | 12 | 5 | 2 |
| pillow | 12 | 5 | 2 |
| plate | 12 | 5 | 0 |
| CLIP B/16 | null | NegCLIP | |
| ADR- , accuracy | 0.7386 [0.714, 0.763] | 0.4779 | 0.7394 |
| global, oracle threshold | 0.5387 | 0.5321 | 0.5287 |
| viewer-frame consistency ( ) | 0.9290 [0.905, 0.951] | 0.5069 | 0.9172 |
| object-frame consistency ( ) | 0.2239 [0.134, 0.328] | 0.4776 | 0.2239 |
| both frames ( ) | 0.8667 | 0.3333 | 0.9667 |
| specificity on false statements ( ) | 0.6281 | 0.4606 | 0.6348 |
| VG-Rel L/R, | What’sUp-A, | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ADR- | ADR- | ADR- | ADR- | |||||||||
| acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | acc. | n.acc. | gain | |
| 90.35 | 49.26 | 88.72 | 50.66 | 29.41 | 25.25 | 60.78 | 25.00 | |||||
| 90.45 | 49.03 | 90.22 | 49.44 | 31.62 | 25.74 | 41.18 | 25.25 | |||||
| 90.40 | 48.90 | 90.25 | 48.80 | 31.86 | 24.75 | 33.82 | 25.00 | |||||
| 90.40 | 48.88 | 90.30 | 48.26 | 32.35 | 24.75 | 32.11 | 25.00 | |||||
| Readout | param. | param. vs. | FLOPs | FLOPs vs. | mem. (GiB) | mem. vs. | time (ms) | time vs. |
|---|---|---|---|---|---|---|---|---|
| CLIP | CLIP | CLIP | CLIP | |||||
| deployed / CLIP | 0.15 B | 1.0 | 1.0 | 0.60 | 1.0 | 9.28 | 1.0 | |
| ADR / CLIP | 0.15 B | 1.0 | 3.3 | 0.66 | 1.1 | 20.78 | 2.2 | |
| half-crop / CLIP | 0.15 B | 1.0 | 4.0 | 0.60 | 1.0 | 43.83 | 4.7 | |
| deployed / SigLIP2 | 1.14 B | 7.6 | 21.4 | 4.35 | 7.3 | 35.97 | 3.9 | |
| ADR / SigLIP2 | 1.14 B | 7.6 | 30.4 | 4.40 | 7.3 | 52.98 | 5.7 |
| ADR- , the default | |
|---|---|
| definition | the excess of the Sinkhorn coupling |
| what it buys | The underlying coupling constrains both marginals, so words share a patch budget. Subtracting the uniform baseline removes its contribution to the centroid. The retained weights no longer preserve either marginal. Their column masses can be inspected alongside the localization maps, but do not by themselves establish localization quality. |
| what it costs | Sinkhorn scaling requires iteration. The shared patch budget in couples word localizations and can limit how well objects of different sizes are represented. A noun can lose all its excess mass and then has no centroid (App. D.3 ). |
| ADR- | |
| definition | the softmax of over patches, computed independently for each word |
| what it buys | Closed form with no iteration. Each word is localized on its own, so a large object cannot drain mass from a small one. The column mass is fixed, so every centroid is defined. |
| Subset | Dataset | Relations | Classes | Question type | Items |
| (a) Direction-balanced subsets (Tabs. 2 and 3 ) | |||||
| VG-Relation left/right | ARO | Horizontal | 2 | Multiple choice | 3918 |
| COCO-spatial | What’sUp | Horizontal, vertical | 2 | Multiple choice | 404 |
| What’sUp-A | What’sUp | Horizontal, vertical | 4 | Multiple choice | 408 |
| What’sUp-B | What’sUp | Horizontal, depth | 2 | Multiple choice | 354 |
| VSR horizontal | VSR | Horizontal | 2 | True/false | 1201 |
| evaluation set | source | selection | items |
|---|---|---|---|
| VG-Relation left/right, 6000-item cell | ARO | relation field, 6000-item pick | 3918 |
| VG-Relation vertical, 6000-item cell | ARO | relation field, 6000-item pick | 800 |
| COCO-spatial two-object | What’sUp | all two-object pairs | 440 (404 covered) |
| What’sUp-A | What’sUp | full set, 4-way | 412 (408 covered) |
| What’sUp-B | What’sUp | front/behind and left/right pairs | 354 |
| VSR horizontal | VSR | horizontal subset | 1224 (1201 covered) |
| model | note |
|---|---|
| (a) frozen dual encoders, deployed pooled score and ADR (Tabs. 2 and 6 ) | |
| CLIP ViT-B/32 | |
| CLIP ViT-B/16 ⋆ | |
| CLIP ViT-L/14 | |
| LAION-2B ViT-B/16 ⋆ | App. A.3 |
| DFN5B ViT-H/14 ⋆ | |