CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when scoring rules treat the subject and object symmetrically, direction cancels regardless of encoder training. Guided by this analysis, we introduce Antisymmetric Displacement Readout (ADR), which aligns caption words with image patches in the frozen features and scores each relation by the signed displacement between matched object centroids. Notably, ADR succeeds without additional training or learned parameters, thereby demonstrating that directional information remains in the frozen encoder. However, text and world priors can inflate accuracy, so we further introduce prior deflation, which measures the benefit of the image-text pairing as the grounded gain over a null that pairs each item with an unrelated image. Extensive experiments across encoder families show that ADR substantially improves over deployed scores, which remain near chance on most direction-balanced sets even for fine-tuned encoders. Compared with more complex readouts, ADR outperforms the evaluated MLLM likelihood readouts and is competitive with their chat inference at a small fraction of the computation. These results support our claim that directional information can be recovered from frozen features by an appropriate readout. Our implementation and evaluation kit will be publicly available.
Figures & tables
Figure 1: Readout blindness and ADR. (a) Choosing between captions with opposite relation directions. The deployed score gives near-equal scores to both, while ADR recovers the direction from the signed centroid displacement D from the same frozen encoder. (b) Grounded gain on VG-Relation left/right. Dots mark the best readout per group, and segments show the range.
Figure 2: ADR (top) versus deployed score (bottom) on the same frozen VLM. ADR (1) computes patch–word similarity S , (2) keeps the excess Sinkhorn coupling π~ , (3) localizes the subject and object centroids, and (4) scores their signed displacement D=er⊤(cs−co) along the relation axis. The deployed score compares pooled features (CLS token for CLIP, attention-pooled token for SigLIP) and gives near-equal scores to a caption and its argument swap, leaving direction ambiguous.
Subset
Random
Metric
CLIP
NegCLIP
SigLIP2
LLaVA-1.5
Qwen3.6
likelihood
likelihood
Prior-based subset
50.03
acc.
72.71
85.14
62.07
93.26
93.99
VALSE actant swap
n.acc.
41.31
62.80
28.35
88.51
88.41
dev.
−8.7
+12.8
−21.7
+38.5
+38.4
gain
+31.4
+22.3
+33.7
+4.8
+5.6
Direction-balanced subset
50.01
acc.
49.64
50.26
49.95
65.80
74.76
Table 1: Prior effects on a prior-biased subset (VALSE actant swap) and a direction-balanced subset (VG-Relation left/right). Acc. is accuracy with the original images. n.acc. is accuracy under the derangement null (Sec. 6 ). Random is accuracy from uniformly random predictions. Dev. ( n.acc.−random ) measures deviation from random guessing. Gain is grounded gain ( acc.−n.acc. ) and measures the benefit of the original image-caption alignment. Accuracies are reported in percent (%), and dev. and gain in percentage points. Full results on all eight subsets are given in Tab. 5 .
Figure 3: Prior deflation compares accuracy under the original pairing and a derangement null. Their difference is grounded gain.
Readout
VG-Rel L/R
COCO
WU-A
WU-B
VSR
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
Random
50.01
50.01
0.0
49.99
49.99
0.0
25.02
25.02
0.0
50.01
50.01
0.0
50.01
50.01
0.0
CLIP
49.64
49.69
−0.1
45.30
47.28
−2.0
26.96
25.98
+1.0
53.67
52.83
+0.8
53.87
53.21
+0.7
ADR- π~
90.45
49.03
+41.4
92.33
51.73
+40.6
31.62
25.74
+5.9
93.22
52.26
+41.0
73.86
47.79
+26.1
ADR- πd
90.22
49.44
+40.8
94.06
52.72
+41.3
41.18
25.25
+15.9
98.02
50.00
+48.0
74.94
47.88
+27.1
LAION-2B
49.13
49.77
−0.6
43.07
47.28
−4.2
27.70
26.23
+1.5
50.00
50.28
−0.3
51.96
52.96
−1.0
Table 2: Random guess, deployed readouts and ADR across nine model families on five direction-balanced sets. Each subset reports accuracy with the original images (acc.), derangement-null accuracy (n.acc.), and grounded gain (gain). Accuracies are in percent (%) and gains in percentage points. Grey rows show deployed readouts, followed by ADR- π~ and ADR- πd on the same encoder and items. Red highlights grounded gain, the main metric for comparing readouts.
Readout
VG-Rel L/R
COCO
WU-A
WU-B
VSR
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
Deployed readouts
CLIP
49.64
49.69
−0.1
45.30
47.28
−2.0
26.96
25.98
+1.0
53.67
52.83
+0.8
53.87
53.21
+0.7
SigLIP2
49.95
50.54
−0.6
49.01
45.05
+4.0
29.17
24.75
+4.4
51.69
48.87
+2.8
54.70
52.96
+1.7
Half-crop
CLIP
83.67
49.16
+34.5
82.92
55.45
+27.5
44.61
22.30
+22.3
100.00
50.00
+50.0
70.52
47.88
+22.6
Table 3: Deployed readouts, alternative readouts and ADR on five direction-balanced sets. Each subset reports accuracy with the original images (acc.), derangement-null accuracy (n.acc.), and grounded gain (gain). Accuracies in percent (%) and gains in percentage points. All readouts use the same evaluation items as Tab. 2 . Red highlights grounded gain. Bold marks the highest gain.
Figure 4: ADR predictions and reliability. (a) Correct predictions for horizontal, vertical and depth relations. Red and blue green maps show subject and object excess mass. Circles mark centroids, and arrows point from object to subject. Margin is the score difference between correct and incorrect captions, D(C+)−D(C−) for ADR. (b) A localization failure of ADR- π~ . Curves show the fraction of each word’s excess mass at each image height. The diffuse mass for "dining" produces an incorrect centroid order and prediction. (c) Accuracy on the binary direction-balanced sets after retaining samples with the largest absolute margins. Increasing the margin threshold retains fewer samples but yields higher ADR accuracy, showing that its margin helps select reliable predictions.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
CLIP ViT-B/16
SigLIP2-So400m
Median text difference
0.2251
0.2494
Pairs with identical text features
0
0
Global score ties
0
0
Global real / null accuracy
49.64 / 49.69
49.95 / 50.54
Global wrong, ADR- π~ correct, item count
1785
1799
Median text difference, global wrong and ADR- π~ correct
0.2264
0.2505
Appendix
Table 4: Global cosine diagnostics on VG-Relation left/right ( n=3918 ). Text differences are Euclidean distances between unit-norm caption features. The last two rows restrict the comparison to items correctly ranked by ADR- π~ . Accuracies are percentages.
Subset
random
metric
CLIP
NegCLIP
SigLIP2
LLaVA-1.5
Qwen3.6
likelihood
likelihood
Direction-balanced subsets
VG-Rel left/right
50.01
acc.
49.64
50.26
49.95
65.80
74.76
n.acc.
49.69
49.67
50.54
49.44
49.29
dev.
−0.3
−0.3
+0.5
−0.6
−0.7
gain
−0.1
+0.6
−0.6
+16.4
+25.5
Appendix
Table 5: Prior effects across five readouts on all eight evaluation subsets. The five direction-balanced subsets are used for the main comparisons, and the three prior-biased subsets are reported separately. Acc., n.acc., dev. and gain follow Tab. 1 . Random is accuracy from uniformly random predictions. Accuracies are in percent (%), and dev. and gain in percentage points.
Readout
VG-Rel L/R
COCO
WU-A
WU-B
VSR
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
n=3918
n=404
n=408
n=354
n=1201
random
50.01
50.01
0.0
49.99
49.99
0.0
25.02
25.02
0.0
50.01
50.01
0.0
50.01
50.01
0.0
frozen dual encoders, deployed score and ADR
CLIP ViT-B/16
49.64
49.69
−0.1
45.30
47.28
−2.0
26.96
25.98
+1.0
53.67
52.83
+0.8
53.87
53.21
+0.7
ADR- π~
90.45
49.03
+41.4
92.33
51.73
+40.6
31.62
25.74
+5.9
93.22
52.26
+41.0
73.86
47.79
+26.1
Appendix
Table 6: Readouts on the five direction-balanced subsets. Each subset reports acc., n.acc., and gain as defined in Tab. 2 . Accuracies are in percent and gains in percentage points. Grey rows show deployed scores, followed by ADR with four choices of alignment weights (App. D.2 ). VSR exceptions are marked and explained below the table. The random-guess row averages 10,000 uniform draws.
readout
VG-Rel L/R
COCO
WU-B
VSR
pooled
n
3918
404
354
1201
5877
(a) AUROC (derangement null)
ADR- π~
0.891 (0.510)
0.904 (0.516)
0.888 (0.449)
0.713 (0.484)
0.863 (0.504)
ADR- πd
0.880 (0.509)
0.899 (0.483)
0.970 (0.506)
0.694 (0.478)
0.844 (0.504)
deployed cosine
0.509 (0.505)
0.461 (0.461)
0.512 (0.481)
0.504 (0.487)
0.514 (0.504)
MLLM likelihood, Qwen3.6-27B
0.717 (0.496)
0.726 (0.555)
0.829 (0.508)
0.568 (0.461)
0.661 (0.498)
Appendix
Table 7: Risk and coverage on the four binary direction-balanced sets. The pooled results use only the four direction-balanced sets ( 5877 items). The dual-encoder readouts use CLIP ViT-B/16, and the likelihood readout uses Qwen3.6-27B. (a) AUROC of the absolute margin as a ranking of correct above wrong items, with the same statistic under the derangement null in parentheses. (b) Accuracy of the most confident 30% of the items. Pooling ranks the items of the four direction-balanced sets by one margin, so it is the accuracy a single threshold would obtain. The four-way What’sUp-A is excluded from this binary-set analysis.
subject noun
items
ViT-B/16
SigLIP2
bottle
28
8
2
bowl
12
6
3
cup
12
6
0
orange
12
5
2
pillow
12
5
2
plate
12
5
0
Appendix
Table 8: Relative-position failures grouped by subject noun on What’sUp-A, with both encoders evaluated on the same items. Rows show the thirteen nouns with at least twelve items, ordered by the ViT-B/16 failure count; the remaining forty-two are pooled. Counts indicate failures of the estimated relative positions, without identifying which object is mislocalized. SigLIP2 has no such failures for 29 of the fifty-five subject-noun groups.
CLIP B/16
null
NegCLIP
ADR- π~ , accuracy
0.7386 [0.714, 0.763]
0.4779
0.7394
global, oracle threshold
0.5387
0.5321
0.5287
viewer-frame consistency ( n=507 )
0.9290 [0.905, 0.951]
0.5069
0.9172
object-frame consistency ( n=67 )
0.2239 [0.134, 0.328]
0.4776
0.2239
both frames ( n=30 )
0.8667
0.3333
0.9667
specificity on false statements ( n=597 )
0.6281
0.4606
0.6348
Appendix
Table 9: VSR covered horizontal subset ( n=1201 , whole images, third-party labels) and its decomposition by the benchmark’s reference-frame annotation. Frame rows are true-item consistency rates, the false-statement row is the shared specificity, and balanced accuracy pairs each frame’s consistency with that specificity.
VG-Rel L/R, n=3918
What’sUp-A, n=408
ε
ADR- π~
ADR- πd
ADR- π~
ADR- πd
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
acc.
n.acc.
gain
0.02
90.35
49.26
+41.1
88.72
50.66
+38.1
29.41
25.25
+4.2
60.78
25.00
+35.8
0.05
90.45
49.03
+41.4
90.22
49.44
+40.8
31.62
25.74
+5.9
41.18
25.25
+15.9
0.1
90.40
48.90
+41.5
90.25
48.80
+41.5
31.86
24.75
+7.1
33.82
25.00
+8.8
0.2
90.40
48.88
+41.5
90.30
48.26
+42.0
32.35
24.75
+7.6
32.11
25.00
+7.1
Appendix
Table 10: Temperature sensitivity on frozen CLIP ViT-B/16 features. Each variant reports acc., n.acc., and gain as defined in Tab. 2 . Accuracies are in percent and gains in percentage points. The default temperature is ε=0.05 .
Readout
param.
param. vs.
FLOPs
FLOPs vs.
mem. (GiB)
mem. vs.
time (ms)
time vs.
CLIP
CLIP
CLIP
CLIP
deployed / CLIP
0.15 B
1.0 ×
3.45×1010
1.0 ×
0.60
1.0 ×
9.28
1.0 ×
ADR / CLIP
0.15 B
1.0 ×
1.15×1011
3.3 ×
0.66
1.1 ×
20.78
2.2 ×
half-crop / CLIP
0.15 B
1.0 ×
1.38×1011
4.0 ×
0.60
1.0 ×
43.83
4.7 ×
deployed / SigLIP2
1.14 B
7.6 ×
7.39×1011
21.4 ×
4.35
7.3 ×
35.97
3.9 ×
ADR / SigLIP2
1.14 B
7.6 ×
1.05×1012
30.4 ×
4.40
7.3 ×
52.98
5.7 ×
Appendix
Table 11: Computational cost per item, averaged over the same 50 items on one NVIDIA H200 after one warm-up call. All ratios are relative to the deployed CLIP.
ADR- π~ , the default
definition
the excess π~=[π0−Nm1]+ of the Sinkhorn coupling
what it buys
The underlying coupling π0 constrains both marginals, so words share a patch budget. Subtracting the uniform baseline removes its contribution to the centroid. The retained weights π~ no longer preserve either marginal. Their column masses can be inspected alongside the localization maps, but do not by themselves establish localization quality.
what it costs
Sinkhorn scaling requires iteration. The shared patch budget in π0 couples word localizations and can limit how well objects of different sizes are represented. A noun can lose all its excess mass and then has no centroid (App. D.3 ).
ADR- πd
definition
the softmax of S/ε over patches, computed independently for each word
what it buys
Closed form with no iteration. Each word is localized on its own, so a large object cannot drain mass from a small one. The column mass is fixed, so every centroid is defined.
Appendix
Table 12: The column weights of ADR. Every weight feeds the same relation axis, sign rule and zero threshold and differs only in how patch mass is assigned to words. Their accuracies are reported in Tab. 6 .
Subset
Dataset
Relations
Classes
Question type
Items
(a) Direction-balanced subsets (Tabs. 2 and 3 )
VG-Relation left/right
ARO
Horizontal
2
Multiple choice
3918
COCO-spatial
What’sUp
Horizontal, vertical
2
Multiple choice
404
What’sUp-A
What’sUp
Horizontal, vertical
4
Multiple choice
408
What’sUp-B
What’sUp
Horizontal, depth
2
Multiple choice
354
VSR horizontal
VSR
Horizontal
2
True/false
1201
Appendix
Table 13: Evaluation subsets from ARO ( Yuksekgonul et al., 2023 ) , What’sUp ( Kamath et al., 2023 ) , VSR ( Liu et al., 2023 ) and VALSE ( Parcalabescu et al., 2022 ) . Classes gives the number of possible answers, and Items counts evaluated examples. Data processing details are given in App. E.1 .
evaluation set
source
selection
items
VG-Relation left/right, 6000-item cell
ARO
relation field, 6000-item pick
3918
VG-Relation vertical, 6000-item cell
ARO
relation field, 6000-item pick
800
COCO-spatial two-object
What’sUp
all two-object pairs
440 (404 covered)
What’sUp-A
What’sUp
full set, 4-way
412 (408 covered)
What’sUp-B
What’sUp
front/behind and left/right pairs
354
VSR horizontal
VSR
horizontal subset
1224 (1201 covered)
Appendix
Table 14: Evaluation sets. Each set is a deterministic selection from a public benchmark. The VG-Relation results use the 6000-item sample. Sources are ARO VG-Relation ( Yuksekgonul et al., 2023 ) , the What’sUp release ( Kamath et al., 2023 ) , which also supplies COCO-spatial, and VSR ( Liu et al., 2023 ) . Covered counts the items whose caption the fixed parsing rule covers, and VSR carries the benchmark’s own reference-frame labels. COCO-spatial’s 2247 single-object captions lie outside the two-argument readout.
model
note
(a) frozen dual encoders, deployed pooled score and ADR (Tabs. 2 and 6 )
CLIP ViT-B/32
CLIP ViT-B/16 ⋆
CLIP ViT-L/14
LAION-2B ViT-B/16 ⋆
App. A.3
DFN5B ViT-H/14 ⋆
Appendix
Table 15: Every model evaluated in this paper. ⋆ marks the nine score families of Table 2 , where BLIP enters through the ITC branch of its base checkpoint. Loader settings are in the released code.
Video Large Language Models (Video-LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed image-plane motion direction. On simple videos of a single object moving left, right, up, or down, most Video-LLMs perform near chance, with above-chance cases largely attributable to prediction biases rather than genuine direction understanding. We call this failure directional motion blindness. We localize the failure by tracing motion direction information through the Video-LLM pipeline. Motion direction remains linearly accessible from the vision encoder, projector, and LLM hidden states, but the readout fails to bind this signal to the correct verbal answer option, revealing a direction binding gap. Although synthetic motion direction instruction tuning reduces this gap on the source domain, motion direction concept vector analysis shows that visual complexity weakens the signal magnitude and limits out-of-domain generalization. We introduce MoDirect, a dataset family for motion direction instruction tuning and evaluation, and DeltaDirect, a diagnosis-driven, projector-level objective that predicts normalized 2-D motion vectors from adjacent-frame feature deltas. On MoDirect-SynBench, instruction tuning with DeltaDirect improves motion direction accuracy from 25.9% to 85.9%. On MODIRECT-REALBENCH, DeltaDirect improves realworld motion direction accuracy by 21.4 points over the vanilla baseline without real-world tuning data, while preserving standard video-understanding performance. Our project page is available at https://jong980812.github.io/which-way-did-it-move/
Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can recognize every word and object yet prefer a compositionally incorrect caption. We argue that a frozen encoder retains this association structure, so the problem is to read it rather than to rebuild it beside the pretrained similarity. We introduce BindCLIP, a pairwise scorer built on one latent object: a balanced token--patch--depth optimal-transport coupling that places both candidate captions and several visual depths in a single plan. Semantic, entity, order, and spatial evidence are read as energies of this state, and exchanging the candidates permutes the plan, making the score exactly antisymmetric. A geometric refinement inside the coupling contracts moves that the candidates and the visual depths do not support. No task label, parser, relation inventory, or detector is used. One checkpoint and one inference path improve the official What'sUp, ARO, and SugarCrepe benchmarks over frozen global CLIP, with the strongest transfer on the relation splits. Controls rule out patch access and caption-length shortcuts, and an inference-time lesion localizes spatial arrangement to the coupling.
Liuyang Song, Yi Zhang, Zhongyi Deng +2
Peking University · Dongguan University of Technology · Sichuan Agricultural University
Vision-language models (VLMs) achieve strong visual reasoning performance, yet subtle changes from routine image capture and processing can alter their reasoning trajectories even when images appear nearly identical. In long-horizon generation, the resulting activation shifts may accumulate across decoding steps, progressively altering reasoning tokens and ultimately changing the final answer, a phenomenon referred to as answer flips. To address this instability, we propose FlipDir (Flip-Direction Steering), a training-free inference-time method that estimates a low-rank flip-inducing activation subspace from contrastive pairs of original and answer-flipping inputs and selectively steers hidden states during decoding. A margin-based gate limits subspace attenuation to uncertain decoding steps, recovering original predictions while preserving stable ones. To evaluate robustness beyond accuracy or consistency on fixed test sets, we introduce VisFlip, a benchmark framework that constructs evaluation groups for a target model and visual variation setting to separately assess recovery of original predictions and preservation of stable ones. VisFlip spans nine dataset-variation combinations across scientific reasoning, robot-scene understanding, and medical VQA, covering subtle visual variations common in each domain. Experiments across 18 settings demonstrate that FlipDir consistently outperforms existing methods on the combined recovery and preservation metric. We will make our code publicly available.
Yeonsung Jung, Joonhyun Jeong, Hoang Pham +4
Graduate School of AI, KAIST · NAVER Cloud · Department of Computer Science and Engineering, The Ohio State University +3