Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning
Authors: Yeonsung Jung, Joonhyun Jeong, Hoang Pham, Joowon Kim, Yoonsik Park, Viet Dac Lai, Eunho Yang
Organizations: Graduate School of AI, KAIST · NAVER Cloud · Department of Computer Science and Engineering, The Ohio State University · Department of Biomedical Informatics, The Ohio State University · Adobe Research · AITRICS
Vision-language models (VLMs) achieve strong visual reasoning performance, yet subtle changes from routine image capture and processing can alter their reasoning trajectories even when images appear nearly identical. In long-horizon generation, the resulting activation shifts may accumulate across decoding steps, progressively altering reasoning tokens and ultimately changing the final answer, a phenomenon referred to as answer flips. To address this instability, we propose FlipDir (Flip-Direction Steering), a training-free inference-time method that estimates a low-rank flip-inducing activation subspace from contrastive pairs of original and answer-flipping inputs and selectively steers hidden states during decoding. A margin-based gate limits subspace attenuation to uncertain decoding steps, recovering original predictions while preserving stable ones. To evaluate robustness beyond accuracy or consistency on fixed test sets, we introduce VisFlip, a benchmark framework that constructs evaluation groups for a target model and visual variation setting to separately assess recovery of original predictions and preservation of stable ones. VisFlip spans nine dataset-variation combinations across scientific reasoning, robot-scene understanding, and medical VQA, covering subtle visual variations common in each domain. Experiments across 18 settings demonstrate that FlipDir consistently outperforms existing methods on the combined recovery and preservation metric. We will make our code publicly available.
Figures & tables
Figure 1: A subtle exposure shift can flip the final answer. For nearly identical Robo2VLM images, Qwen3-VL-8B produces identical text before diverging at a versus another and reaching different action predictions. Token probabilities are schematic.
GPT-5.5
Claude Opus 4.8
Qwen3-VL-8B
Gemma-3-12B
Clean accuracy (%)
93.2
91.8
68.3
47.8
Answer flip (%)
9.6
21.7
47.6
62.4
Table 1: Answer flips on recent models.
Figure 3
Domain
Source
Train
Val
Test
Coverage
Science reasoning
M3CoT
1,582
221
450
Forces & motion, energy, heat, magnetism
Robot scenes
Robo2VLM
1,500
200
567
Robot state, reachability, depth, next action
Medical VQA
GMAI-MMBench
300
200
200
X-ray, CT, endoscopy, fundus
Table 2: Composition of VisFlip across three domains.
Figure 4: Recovery–preservation trade-offs in selected settings. Each point is a reported configuration, not a tuning sweep. All 18 settings, including Base, appear in Appendix E.1 .
Figure 6Table 7
Base
LEAD
VTI
VCD
FlipDir
Offline directions
No
No
Yes
No
Yes
Decode FLOPs
1.00×
≈1.08×
≈1.00×
≈2.00×
≈1.08×
Table 5: Offline preparation and mean estimated decoding FLOPs relative to Base.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Pair found
Pair not found
Domain
Searches
Calls/search
Calls
Searches
Calls/search
Calls
All calls
Gemma-3-12B
M3CoT
3,817
6.25±5.08
23,858
2,276
16.82±5.28
38,271
62,129
Robo2VLM
1,485
6.99±5.09
10,387
4,716
16.10±1.90
75,946
86,333
GMAI-MMBench
781
6.43±5.00
5,021
719
17.02±5.90
12,239
17,260
Qwen3-VL-8B
Appendix
Table 6: VLM calls for benchmark construction. Calls per search are reported as mean ± standard deviation. Pair found denotes obtaining both a flip and a non-flip child.
Table 7: Perturbation ranges. Each instance uses a deterministic seed chain. Appendix A.5 specifies modality-dependent medical operators.
Qwen3-VL-8B
Gemma-3-12B
Setting
U pairs
Eval C/N/F
U pairs
Eval C/N/F
M3CoT exposure
737
429/428/183
989
449/447/278
M3CoT white balance
723
429/428/190
977
449/448/272
M3CoT JPEG
759
429/428/187
1039
449/444/298
Robo2VLM exposure
334
560/558/107
316
567/566/149
Robo2VLM lens smudge
255
560/554/86
319
567/566/153
Appendix
Table 8: Subspace and evaluation pool sizes. U pairs counts the pairs used to estimate the subspace, at most one per parent. Eval C/N/F lists the actual clean, non-flip, and flip denominators used for recovery metrics. Parent pools are given in Table 2 .
Setting
Qwen3-VL-8B
Gemma-3-12B
M3CoT exposure
42.7% (n=429)
61.9% (n=449)
M3CoT white balance
44.3% (n=429)
60.6% (n=449)
M3CoT JPEG
43.6% (n=429)
66.4% (n=449)
Robo2VLM exposure
19.1% (n=560)
26.3% (n=567)
Robo2VLM lens smudge
15.4% (n=560)
27.0% (n=567)
Robo2VLM motion blur
16.8% (n=560)
33.0% (n=567)
Appendix
Table 9: Answer flips in the evaluation pools. Percentages are 100nF/nC , using the flip and clean counts in Table 8 ; n=nC .
Gemma-3-12B
Qwen3-VL-8B
Setting
ℓ
α
H(C,N,F)
ℓ
α
H(C,N,F)
M3CoT exposure
12
1.0
61.6
4
1.0
78.3
M3CoT white balance
12
0.5
60.8
36
1.0
76.4
M3CoT JPEG
12
1.0
63.0
36
1.0
78.9
Robo2VLM exposure
12
0.5
73.4
36
1.0
82.4
Robo2VLM lens smudge
36
0.5
73.6
36
1.0
82.5
Appendix
Table 10: Selected configurations. Layer ℓ , strength α , and main-table H(C,N,F) . All settings use k=8 , gate thresholds (0.3,0.7) , and uncentered SVD with at most one flip pair per parent.
Figure 8: Additional activation and decoding analyses. (a) Cumulative energy at layer 4 for the same 255 Qwen3-VL-8B/Robo2VLM lens-smudge pairs as Figure 3 . (b) Divergence and non-divergence token margins for Gemma-3-12B, complementing the Qwen results in Figure 3 .
Figure 9: Recovery–preservation trade-offs for Qwen3-VL-8B. All nine settings and five methods are shown on shared axes. Higher and further right indicate better recovery and preservation, respectively.
Figure 10: Recovery–preservation trade-offs for Gemma-3-12B. All nine settings, with the same methods and axes as Figure 9 . Base is at (100,0) in every setting.
Model
Target setting
Source U
H(C,N,F)
Target-specific
Qwen3-VL-8B
Robo2VLM lens
M3CoT exposure
75.7
82.5
Qwen3-VL-8B
Robo2VLM lens
GMAI detector
76.3
82.5
Qwen3-VL-8B
M3CoT exposure
Robo2VLM lens
71.5
78.3
Qwen3-VL-8B
M3CoT exposure
GMAI detector
74.2
78.3
Qwen3-VL-8B
GMAI detector
M3CoT exposure
65.3
70.7
Qwen3-VL-8B
GMAI detector
Robo2VLM lens
69.7
70.7
Appendix
Table 11: Transfer results. H(C,N,F) using the source subspace and the target-specific layer and strength. The last column uses a target-matched subspace.
Model
Setting
k=2
4
8
16
32
Qwen3-VL-8B
M3CoT exposure
71.0
74.3
78.3
73.2
72.6
Robo2VLM lens smudge
82.2
79.6
82.5
79.1
77.8
GMAI rotation
65.4
64.7
70.7
63.7
69.6
Gemma-3-12B
M3CoT exposure
59.2
58.6
61.6
59.2
60.2
Robo2VLM lens smudge
77.1
70.4
73.6
71.4
70.2
GMAI rotation
65.1
64.2
64.2
65.4
64.5
Appendix
Table 12: Rank sensitivity. H(C,N,F) with all other components fixed to Table 10 . The shared rank k=8 is shaded; row maxima are bold.
Gate Off
Gate On
Setting
C
N
F
H
C
N
F
H
ΔH
Gemma-3-12B
M3CoT exposure
59.5
63.8
45.0
54.8
70.2
69.1
50.0
61.6
+6.8
Robo2VLM lens smudge
81.7
83.4
51.6
68.8
87.8
86.4
56.2
73.6
+4.8
GMAI rotation
84.0
78.7
37.3
58.3
84.5
83.2
43.7
64.2
+5.9
Qwen3-VL-8B
Appendix
Table 13: Margin-gate ablation. C/N/F/H (%) with and without the gate. ΔH=HOn−HOff , computed from the displayed values.
Figure 11: Performance gains from margin gating. Changes in clean preservation (C), non-flip preservation (N), flip recovery (F), and their harmonic mean (H), measured in percentage points as Gate On minus Gate Off. All panels share the same scale; values are computed from Table 13 . The subspace, layer, and strength are fixed within each comparison. GMAI denotes GMAI-MMBench.
Setting
50 pairs
subsampled
full pool
Gemma M3CoT exposure
60.2
56.8 (250)
61.6 (989)
Gemma Robo2VLM lens
67.0
72.2 (150)
73.6 (319)
Gemma GMAI detector
67.0
63.8 (100)
64.2 (192)
Qwen Robo2VLM lens
78.6
79.2 (150)
82.5 (255)
Appendix
Table 14: Subspace-estimation sample size. H(C,N,F) with other components fixed; parent counts are shown in parentheses.
Gate
Seconds per parent
Setting
Open %
Entries
FlipDir
VCD
LEAD
VTI
Qwen M3CoT exposure
15.5
9.1
254.1
260.5
92.4
252.2
Qwen M3CoT wb
14.5
8.5
231.3
265.4
135.1
227.2
Qwen M3CoT JPEG
12.9
8.0
232.7
258.7
134.6
241.4
Qwen Robo2VLM exp.
17.3
9.6
121.8
109.0
71.9
44.7
Qwen Robo2VLM lens
N/A
N/A
71.8
96.0
49.2
43.7
Appendix
Table 15: Runtime and gate statistics. Time is measured in seconds per parent under the shared serving regime. Gate statistics are replayed on stored unedited traces; entries are counted per 100 tokens. N/A denotes unavailable measurements.
Figure 12: A repaired flip on Gemma-3-12B Robo2VLM exposure shift. The unedited pass on the varied image restates the completed phase and flips to (A); with FlipDir the reasoning advances to the cycle’s next phase and the answer returns to (B), the clean answer and the ground truth. Colored spans mark where the trajectories diverge.
Figure 13: M3CoT variation intensities on the example of Figure , rendered from the fixed bands. Rows: exposure shift at the EV band endpoints ( +0.45 , +0.65 ), white-balance shift at +1000 K and +3000 K, and the mildest ( q=95,85 ) and strongest ( q=85,72 ) double-JPEG draws.
Figure 14: Robo2VLM lens-smudge example from VisFlip . The perturbation composites wipe smears, droplets, dust, local blur, and mild haze while preserving the underlying scene.
Figure 15: GMAI-MMBench variation intensities on an endoscopy image. Rows: window-level shift at the joint contrast and gamma endpoints, modality-gated defocus blur (disk PSF for this lens modality), and detector rotation at 1.5∘ and 4.0∘ .
Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, directly applying RL to rollout multimodal reasoning can lead to instability, due to the exploitation of language priors, the neglect of visual evidence, and the generation of reasoning traces that are fluent yet not visually grounded. The question arises: Can initially steer the policy toward visually faithful reasoning regime before applying reinforcement learning? To this end, we propose a Faithful Warm-Start (FWS) strategy that first curates samples with explicit vision-language causal relationships from six general VQA benchmarks to construct the FaithfulQA dataset, where each of the image-question pairs gains a certain degree of visual observations, question requirements, commonsense knowledge, domain knowledge, and the final answer. Subsequently, a VLM-based judge is employed to further purify the dataset, ensuring strong causal consistency and visual faithfulness. This warm-start stage equips the model with the capability to understand causally grounded vision-language patterns before subsequent RL optimization under sparse answer-level rewards. Experimental results show that such faithful supervision improves answer accuracy, stabilizes RL training, and reduces visually unsupported reasoning.
Peng, Lee, Yin Zhang +9
GMLab · Hong Kong Polytechnic University · XPENG Robotics
Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising solution by optimizing policies using answer correctness signals. Despite their effectiveness, prevailing RLVR methods face two critical limitations. First, much of the sampling budget is wasted on trajectories doomed to fail due to early visual description errors. Second, sparse rewards cannot distinguish whether failures stem from visual perception or reasoning stages. We introduce MIRL, a decoupled framework that addresses both limitations by leveraging mutual information (MI) between generated descriptions and visual inputs as a cheap pre-screening signal. This enables intelligent budget allocation toward high-potential trajectories via forking, while decoupled training provides independent MI-based rewards for visual perception optimization, resolving reward blindness. Experiments on six vision-language reasoning benchmarks demonstrate that MIRL achieves 70.22% average accuracy and successfully surpasses the performance of sampling 16 complete trajectories using only 10 pre-samples with top-6 selection (25% fewer complete trajectories). Our code is available at: https://anonymous.4open.science/r/mirl-main/.
Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4
When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual evidence? Correct answers can coexist with flawed reasoning, making accuracy alone an incomplete test of grounding. We introduce VisualFLIP, a paired benchmark with 1,374 images arranged as same-question perturbation pairs across cardinality, attribute, spatial, and logic tasks. Each pair keeps the question fixed but minimally changes the evidence so the gold answer deterministically flips. We evaluate 24 MLLMs with pair accuracy, which requires solving both sides of a pair, and Collapse Rate (CR), which measures how often a model that solves at least one side repeats the same non-empty answer for both images. Together, these metrics show that paired correctness and evidence dependence are related but distinct: capable models can still fail to update after task-critical visual changes, and collapse becomes more severe for some models when the edited image follows an earlier answer in a sequential setting. Further details are available on our project page: https://didizhu-judy.github.io/VisualFLIP/