Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that localizes critical layers and executes visual interventions during the transition from evidence aggregation to semantic decoding. At these layers, a Geometry-Constrained Local Relation Reconstruction (GCLR) module selects and weights vertex-centered visual evidence, while an Adaptive Visual Steering Operator (AVSO) redistributes attention mass toward the selected tokens. Experiments on PGPS9K and PGDP5K show that CVIF raises Overall F1 from 77.85 to 85.58 and from 75.23 to 82.84, respectively, establishing a novel inference-time visual intervention paradigm.
Figures & tables
Figure 1: Attention failure modes of MLLMs in geometric diagram perception. (a) Perceptual errors during the process of geometric reasoning lead to prediction failures. (b) When predicting which line segment intersects BD at point E , shallow layers attend to the visually correct point C . As depth increases, attention shifts toward evidence that confirms a prior-compatible candidate (e.g., AF ); by a later layer, the visually correct point C receives little attention. (c) When determining which angle is annotated by 43∘ , late-layer attention may concentrate on the label “B” and nearby symbols while failing to recover the local topology of the three edges incident to B. In particular, the model does not sufficiently use edge BD as structural evidence for resolving the angle assignment.
Figure 2: Prior-induced local minima and vision-guided trajectory correction. (a) When examining the angle at vertex B, textual priors tend to focus on the conventional ∠ ABC, whereas visual perception is naturally drawn to the salient feature " 43∘ ". As depicted in the energy landscape, this causes the prediction state to be pulled by prior forces (black arrow), making it highly susceptible to being trapped in a neighboring local minimum that represents an erroneous prediction (i.e., ∠ABC=43∘ ). During the critical state, introducing the local topological structure at vertex B provides a vision-driven steering force (blue arrow), thereby reducing the confidence of local minima induced by prior and steering the prediction toward the global minimum. (b) Layer-wise normalized predictive entropy for Qwen3-VL-32B at the question-end token.
Figure 3: The Geometry-Constrained Local Relation Reconstruction (GCLR) mechanism. Using vertex B as a spatial anchor, GCLR identifies its incident edges ( l1 , l2 , l3 ) and nearby geometric annotations. The resulting target token set and weight matrix explicitly compensate for the relative topological relations of the local structures.
Figure 4: Overview of CVIF. Layer-wise normalized predictive entropy provides a diagnostic for selecting candidate critical layers during prediction formation. At these layers, GCLR utilizes geometric vertices as anchors to explicitly decouple and reconstruct spatial-semantic mappings for local structures. Subsequently, the AVSO performs distribution shaping and boost control over the attention, achieving active visual recall and prediction rectification.
Figure 5: The principle of the Adaptive Visual Steering Operator (AVSO).
Methods
Structure
Semantic
Overall
P
R
F1
P
R
F1
P
R
F1
Open-source MLLMs
Qwen2.5-VL-7B-Instruct
73.58
38.00
50.12
38.32
43.82
40.89
47.19
40.70
43.70
InternVL3.5-8B-Instruct
66.14
43.70
52.63
43.42
51.55
47.14
51.44
47.38
49.33
InternVL3.5-30B-Instruct
70.03
58.89
63.98
57.39
56.12
56.75
62.32
57.61
59.87
InternVL3.5-8B
68.90
54.26
60.71
61.52
59.06
60.26
64.00
56.49
60.01
Table 1: Performance comparison on PGPS9K.
Methods
Structure
Semantic
Overall
P
R
F1
P
R
F1
P
R
F1
Qwen2.5-VL-32B-Instruct
70.15
53.01
60.39
63.33
58.9
61.03
65.49
55.74
60.22
Qwen3-VL-8B-Instruct
73.16
51.04
60.13
71.15
60.15
65.19
69.89
55.15
61.65
GLM-4.6V-Flash
76.95
53.18
62.89
71.27
56.73
63.18
72.04
54.78
62.24
Qwen3.5-9B
79.72
65.21
71.74
77.98
72.45
75.11
76.70
68.48
72.36
Qwen3-VL-32B-Instruct
81.28
74.58
77.79
77.43
69.81
73.43
78.26
72.43
75.23
Table 2: Performance comparison on PGDP5K.
Methods
PGPS9K
PGDP5K
Lines
Circles
Structures
Semantics
PPR
Lines
Circles
Structures
Semantics
PPR
Qwen2.5-VL-32B-Instruct
36.57
77.51
34.75
38.42
21.43
34.64
78.64
32.60
38.42
19.36
Qwen3-VL-8B-Instruct
42.79
77.51
38.08
50.75
25.06
42.45
77.85
37.40
50.00
24.00
GLM-4.6V-Flash
43.70
80.60
39.09
49.81
25.77
46.34
80.00
39.70
51.02
25.40
Qwen3.5-9B
56.00
85.43
50.64
61.89
37.36
55.55
82.60
48.30
63.47
36.34
Qwen3-VL-32B-Instruct
64.23
92.57
62.38
58.79
43.47
64.38
87.62
58.08
58.79
40.11
Table 3: Sample Accuracy and Perfect Parsing Rate (PPR) on PGPS9K and PGDP5K.
Methods
Structure
Semantic
Overall
P
R
F1
P
R
F1
P
R
F1
Qwen3-VL-32B (Baseline)
82.91
76.16
79.39
79.63
74.67
77.07
80.37
75.48
77.85
Qwen3-VL-32B 2-stage
84.29
78.92
81.52
81.11
77.85
79.45
81.11
77.85
79.45
Uniform Intv.
86.00
80.13
82.96
79.66
77.03
78.32
81.07
78.70
79.86
w/o GCLR
86.46
81.21
83.75
81.49
78.79
80.12
82.51
80.10
81.28
w/o AVSO
86.12
81.11
83.54
80.34
78.54
79.43
81.12
79.92
80.52
Table 4: Ablation results on PGPS9K.
Figure 6: Qualitative case analysis on representative geometric diagrams.
Methods
Qwen3-VL-32B
Qwen3-VL-8B
InternVL3.5-8B
Qwen2.5-VL-7B
Qwen3-VL-4B
Direct Answer
53.73
45.10
41.18
42.16
41.18
+ Formal Language
53.73
49.61
45.29
46.08
41.37
+ Text Description
58.63
48.24
47.06
46.47
42.94
Table 5: Downstream reasoning accuracy on MathVerse.
Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this work, we propose Procedurally Generated Tasks (PGT), a simple data-driven framework that serves a dual purpose: inducing fine-grained visual understanding and acting as a low-cost diagnostic tool to identify the source of perception failures. By overlaying unambiguous geometric primitives on images, PGT generate additional dense supervision that disentangles visual grounding capability from semantic priors. Extensive experiments on relational, quantitative, and 3D/depth understanding benchmarks show that PGT yields remarkable gains across diverse architectures. Instruction tuning MLLMs on LLaVA-v1.5-Instruct augmented with PGT data results in improvements of up to +20% on the What'sUp benchmark and +13.3% on CV-Bench-2D, while maintaining general perception capabilities. Moreover, finetuning state-of-the-art MLLMs on PGT data leads to boosts of up to +5.5% on What'sUp and +8.3% on CV-Bench-2D. These findings demonstrate that PGT effectively address the bottleneck of fine-grained perception, revealing that many spatial reasoning deficits stem from inadequate supervision signals rather than inherent architectural or resolution limitations.
Rim Assouel, Amir Bar, Michal Drozdzal +1
1Mila - Québec AI Institute · 4McGill University · 3FAIR at Meta Superintelligence Labs +2
Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore these failures separately using two diagnostic paradigms: (1) probing whether visual information is available, via image reconstruction, and (2) measuring multimodal context sensitivity, the extent to which the model follows visual context versus the language prior. To support the second, we introduce the WhatIfVis, a benchmark spanning five coarse-grained dimensions (spatial-temporal, color, count, size, and weight) whose questions admit answers from either the image or the prior. Our analysis yields three findings: (i) Coarse-grained visual evidence is preserved, as these attributes can be reconstructed from the final-layer image tokens of frozen MLLMs. Failures on questions about these attributes therefore point to post-perceptual utilization, rather than to degraded visual encoding during perception. (ii) Even when explicitly instructed to use or ignore visual evidence, vanilla models (without supervised fine-tuning on the WhatIfVis) show unstable visual context sensitivity. Supervised fine-tuning (SFT) improves this controllability and generalizes across domains, and activation patching further localizes the vision-versus-prior trade-off at architecture-specific depths across all six models. (iii) The vision-versus-prior trade-off is controllable along a learned vector. Applying this steering vector, even without any intent instruction, improves controllability over the vanilla model. Together, these results relocate the bottleneck, indicating that for the coarse attributes we study, MLLMs encode the visual evidence but cannot reliably control their reliance on it.
Jiaang Li, Chengzu Li, Zhaochong An +4
University of Copenhagen · University of Cambridge · 3ETH Zürich +1
In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information. The dominant connector-based paradigm projects visual features into textual sequence, enabling unified multimodal alignment and reasoning within a generative architecture. However, our experiments reveal two key limitations: (1) Although visual information serves as the core evidential modality in MLLMs, it is treated on par with textual tokens, diminishing the unique contribution of the visual modality; (2) As generation length increases, particularly within a limited context window, the model's dependence on visual information progressively weakens, resulting in deteriorated vision-language alignment and reduced consistency between generated content and visual semantics. To address these challenges, we propose the Vision Inference Former (VIF), a lightweight architectural module that establishes a direct bridge between pure visual representations and the model's output space. Specifically, VIF continuously injects visual semantics throughout the decoding phase of the inference process, ensuring that the model remains firmly grounded in visual content during generation. We conduct experiments on 14 benchmark tasks covering general reasoning, OCR, table understanding, vision-centric evaluation, and hallucination. Experimental results show that VIF consistently improves model performance across diverse architectures while introducing minimal additional overhead. The code for this work is available at https://github.com/Dong-Xinpeng/VIF.
Xinpeng Dong, Min Zhang, Kairong Han +3
Zhejiang University · East China Normal University · Zhejiang University of Science and Technology