Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning
Authors: Yeonsung Jung, Joonhyun Jeong, Hoang Pham, Joowon Kim, Yoonsik Park, Viet Dac Lai, Eunho Yang
Organizations: Graduate School of AI, KAIST · NAVER Cloud · Department of Computer Science and Engineering, The Ohio State University · Department of Biomedical Informatics, The Ohio State University · Adobe Research · AITRICS
Vision-language models (VLMs) achieve strong visual reasoning performance, yet subtle changes from routine image capture and processing can alter their reasoning trajectories even when images appear nearly identical. In long-horizon generation, the resulting activation shifts may accumulate across decoding steps, progressively altering reasoning tokens and ultimately changing the final answer, a phenomenon referred to as answer flips. To address this instability, we propose FlipDir (Flip-Direction Steering), a training-free inference-time method that estimates a low-rank flip-inducing activation subspace from contrastive pairs of original and answer-flipping inputs and selectively steers hidden states during decoding. A margin-based gate limits subspace attenuation to uncertain decoding steps, recovering original predictions while preserving stable ones. To evaluate robustness beyond accuracy or consistency on fixed test sets, we introduce VisFlip, a benchmark framework that constructs evaluation groups for a target model and visual variation setting to separately assess recovery of original predictions and preservation of stable ones. VisFlip spans nine dataset-variation combinations across scientific reasoning, robot-scene understanding, and medical VQA, covering subtle visual variations common in each domain. Experiments across 18 settings demonstrate that FlipDir consistently outperforms existing methods on the combined recovery and preservation metric. We will make our code publicly available.
Figures & tables
Figure 1: A subtle exposure shift can flip the final answer. For nearly identical Robo2VLM images, Qwen3-VL-8B produces identical text before diverging at a versus another and reaching different action predictions. Token probabilities are schematic.
GPT-5.5
Claude Opus 4.8
Qwen3-VL-8B
Gemma-3-12B
Clean accuracy (%)
93.2
91.8
68.3
47.8
Answer flip (%)
9.6
21.7
47.6
62.4
Table 1: Answer flips on recent models.
Figure 3
Domain
Source
Train
Val
Test
Coverage
Science reasoning
M3CoT
1,582
221
450
Forces & motion, energy, heat, magnetism
Robot scenes
Robo2VLM
1,500
200
567
Robot state, reachability, depth, next action
Medical VQA
GMAI-MMBench
300
200
200
X-ray, CT, endoscopy, fundus
Table 2: Composition of VisFlip across three domains.
Figure 4: Recovery–preservation trade-offs in selected settings. Each point is a reported configuration, not a tuning sweep. All 18 settings, including Base, appear in Appendix E.1 .
Figure 6Table 7
Base
LEAD
VTI
VCD
FlipDir
Offline directions
No
No
Yes
No
Yes
Decode FLOPs
1.00×
≈1.08×
≈1.00×
≈2.00×
≈1.08×
Table 5: Offline preparation and mean estimated decoding FLOPs relative to Base.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Pair found
Pair not found
Domain
Searches
Calls/search
Calls
Searches
Calls/search
Calls
All calls
Gemma-3-12B
M3CoT
3,817
6.25±5.08
23,858
2,276
16.82±5.28
38,271
62,129
Robo2VLM
1,485
6.99±5.09
10,387
4,716
16.10±1.90
75,946
86,333
GMAI-MMBench
781
6.43±5.00
5,021
719
17.02±5.90
12,239
17,260
Qwen3-VL-8B
Appendix
Table 6: VLM calls for benchmark construction. Calls per search are reported as mean ± standard deviation. Pair found denotes obtaining both a flip and a non-flip child.
Table 7: Perturbation ranges. Each instance uses a deterministic seed chain. Appendix A.5 specifies modality-dependent medical operators.
Qwen3-VL-8B
Gemma-3-12B
Setting
U pairs
Eval C/N/F
U pairs
Eval C/N/F
M3CoT exposure
737
429/428/183
989
449/447/278
M3CoT white balance
723
429/428/190
977
449/448/272
M3CoT JPEG
759
429/428/187
1039
449/444/298
Robo2VLM exposure
334
560/558/107
316
567/566/149
Robo2VLM lens smudge
255
560/554/86
319
567/566/153
Appendix
Table 8: Subspace and evaluation pool sizes. U pairs counts the pairs used to estimate the subspace, at most one per parent. Eval C/N/F lists the actual clean, non-flip, and flip denominators used for recovery metrics. Parent pools are given in Table 2 .
Setting
Qwen3-VL-8B
Gemma-3-12B
M3CoT exposure
42.7% (n=429)
61.9% (n=449)
M3CoT white balance
44.3% (n=429)
60.6% (n=449)
M3CoT JPEG
43.6% (n=429)
66.4% (n=449)
Robo2VLM exposure
19.1% (n=560)
26.3% (n=567)
Robo2VLM lens smudge
15.4% (n=560)
27.0% (n=567)
Robo2VLM motion blur
16.8% (n=560)
33.0% (n=567)
Appendix
Table 9: Answer flips in the evaluation pools. Percentages are 100nF/nC , using the flip and clean counts in Table 8 ; n=nC .
Gemma-3-12B
Qwen3-VL-8B
Setting
ℓ
α
H(C,N,F)
ℓ
α
H(C,N,F)
M3CoT exposure
12
1.0
61.6
4
1.0
78.3
M3CoT white balance
12
0.5
60.8
36
1.0
76.4
M3CoT JPEG
12
1.0
63.0
36
1.0
78.9
Robo2VLM exposure
12
0.5
73.4
36
1.0
82.4
Robo2VLM lens smudge
36
0.5
73.6
36
1.0
82.5
Appendix
Table 10: Selected configurations. Layer ℓ , strength α , and main-table H(C,N,F) . All settings use k=8 , gate thresholds (0.3,0.7) , and uncentered SVD with at most one flip pair per parent.
Figure 8: Additional activation and decoding analyses. (a) Cumulative energy at layer 4 for the same 255 Qwen3-VL-8B/Robo2VLM lens-smudge pairs as Figure 3 . (b) Divergence and non-divergence token margins for Gemma-3-12B, complementing the Qwen results in Figure 3 .
Figure 9: Recovery–preservation trade-offs for Qwen3-VL-8B. All nine settings and five methods are shown on shared axes. Higher and further right indicate better recovery and preservation, respectively.
Figure 10: Recovery–preservation trade-offs for Gemma-3-12B. All nine settings, with the same methods and axes as Figure 9 . Base is at (100,0) in every setting.
Model
Target setting
Source U
H(C,N,F)
Target-specific
Qwen3-VL-8B
Robo2VLM lens
M3CoT exposure
75.7
82.5
Qwen3-VL-8B
Robo2VLM lens
GMAI detector
76.3
82.5
Qwen3-VL-8B
M3CoT exposure
Robo2VLM lens
71.5
78.3
Qwen3-VL-8B
M3CoT exposure
GMAI detector
74.2
78.3
Qwen3-VL-8B
GMAI detector
M3CoT exposure
65.3
70.7
Qwen3-VL-8B
GMAI detector
Robo2VLM lens
69.7
70.7
Appendix
Table 11: Transfer results. H(C,N,F) using the source subspace and the target-specific layer and strength. The last column uses a target-matched subspace.
Model
Setting
k=2
4
8
16
32
Qwen3-VL-8B
M3CoT exposure
71.0
74.3
78.3
73.2
72.6
Robo2VLM lens smudge
82.2
79.6
82.5
79.1
77.8
GMAI rotation
65.4
64.7
70.7
63.7
69.6
Gemma-3-12B
M3CoT exposure
59.2
58.6
61.6
59.2
60.2
Robo2VLM lens smudge
77.1
70.4
73.6
71.4
70.2
GMAI rotation
65.1
64.2
64.2
65.4
64.5
Appendix
Table 12: Rank sensitivity. H(C,N,F) with all other components fixed to Table 10 . The shared rank k=8 is shaded; row maxima are bold.
Gate Off
Gate On
Setting
C
N
F
H
C
N
F
H
ΔH
Gemma-3-12B
M3CoT exposure
59.5
63.8
45.0
54.8
70.2
69.1
50.0
61.6
+6.8
Robo2VLM lens smudge
81.7
83.4
51.6
68.8
87.8
86.4
56.2
73.6
+4.8
GMAI rotation
84.0
78.7
37.3
58.3
84.5
83.2
43.7
64.2
+5.9
Qwen3-VL-8B
Appendix
Table 13: Margin-gate ablation. C/N/F/H (%) with and without the gate. ΔH=HOn−HOff , computed from the displayed values.
Figure 11: Performance gains from margin gating. Changes in clean preservation (C), non-flip preservation (N), flip recovery (F), and their harmonic mean (H), measured in percentage points as Gate On minus Gate Off. All panels share the same scale; values are computed from Table 13 . The subspace, layer, and strength are fixed within each comparison. GMAI denotes GMAI-MMBench.
Setting
50 pairs
subsampled
full pool
Gemma M3CoT exposure
60.2
56.8 (250)
61.6 (989)
Gemma Robo2VLM lens
67.0
72.2 (150)
73.6 (319)
Gemma GMAI detector
67.0
63.8 (100)
64.2 (192)
Qwen Robo2VLM lens
78.6
79.2 (150)
82.5 (255)
Appendix
Table 14: Subspace-estimation sample size. H(C,N,F) with other components fixed; parent counts are shown in parentheses.
Gate
Seconds per parent
Setting
Open %
Entries
FlipDir
VCD
LEAD
VTI
Qwen M3CoT exposure
15.5
9.1
254.1
260.5
92.4
252.2
Qwen M3CoT wb
14.5
8.5
231.3
265.4
135.1
227.2
Qwen M3CoT JPEG
12.9
8.0
232.7
258.7
134.6
241.4
Qwen Robo2VLM exp.
17.3
9.6
121.8
109.0
71.9
44.7
Qwen Robo2VLM lens
N/A
N/A
71.8
96.0
49.2
43.7
Appendix
Table 15: Runtime and gate statistics. Time is measured in seconds per parent under the shared serving regime. Gate statistics are replayed on stored unedited traces; entries are counted per 100 tokens. N/A denotes unavailable measurements.
Figure 12: A repaired flip on Gemma-3-12B Robo2VLM exposure shift. The unedited pass on the varied image restates the completed phase and flips to (A); with FlipDir the reasoning advances to the cycle’s next phase and the answer returns to (B), the clean answer and the ground truth. Colored spans mark where the trajectories diverge.
Figure 13: M3CoT variation intensities on the example of Figure , rendered from the fixed bands. Rows: exposure shift at the EV band endpoints ( +0.45 , +0.65 ), white-balance shift at +1000 K and +3000 K, and the mildest ( q=95,85 ) and strongest ( q=85,72 ) double-JPEG draws.
Figure 14: Robo2VLM lens-smudge example from VisFlip . The perturbation composites wipe smears, droplets, dust, local blur, and mild haze while preserving the underlying scene.
Figure 15: GMAI-MMBench variation intensities on an endoscopy image. Rows: window-level shift at the joint contrast and gamma endpoints, modality-gated defocus blur (disk PSF for this lens modality), and detector rotation at 1.5∘ and 4.0∘ .
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4