Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that localizes critical layers and executes visual interventions during the transition from evidence aggregation to semantic decoding. At these layers, a Geometry-Constrained Local Relation Reconstruction (GCLR) module selects and weights vertex-centered visual evidence, while an Adaptive Visual Steering Operator (AVSO) redistributes attention mass toward the selected tokens. Experiments on PGPS9K and PGDP5K show that CVIF raises Overall F1 from 77.85 to 85.58 and from 75.23 to 82.84, respectively, establishing a novel inference-time visual intervention paradigm.
Figures & tables
Figure 1: Attention failure modes of MLLMs in geometric diagram perception. (a) Perceptual errors during the process of geometric reasoning lead to prediction failures. (b) When predicting which line segment intersects BD at point E , shallow layers attend to the visually correct point C . As depth increases, attention shifts toward evidence that confirms a prior-compatible candidate (e.g., AF ); by a later layer, the visually correct point C receives little attention. (c) When determining which angle is annotated by 43∘ , late-layer attention may concentrate on the label “B” and nearby symbols while failing to recover the local topology of the three edges incident to B. In particular, the model does not sufficiently use edge BD as structural evidence for resolving the angle assignment.
Figure 2: Prior-induced local minima and vision-guided trajectory correction. (a) When examining the angle at vertex B, textual priors tend to focus on the conventional ∠ ABC, whereas visual perception is naturally drawn to the salient feature " 43∘ ". As depicted in the energy landscape, this causes the prediction state to be pulled by prior forces (black arrow), making it highly susceptible to being trapped in a neighboring local minimum that represents an erroneous prediction (i.e., ∠ABC=43∘ ). During the critical state, introducing the local topological structure at vertex B provides a vision-driven steering force (blue arrow), thereby reducing the confidence of local minima induced by prior and steering the prediction toward the global minimum. (b) Layer-wise normalized predictive entropy for Qwen3-VL-32B at the question-end token.
Figure 3: The Geometry-Constrained Local Relation Reconstruction (GCLR) mechanism. Using vertex B as a spatial anchor, GCLR identifies its incident edges ( l1 , l2 , l3 ) and nearby geometric annotations. The resulting target token set and weight matrix explicitly compensate for the relative topological relations of the local structures.
Figure 4: Overview of CVIF. Layer-wise normalized predictive entropy provides a diagnostic for selecting candidate critical layers during prediction formation. At these layers, GCLR utilizes geometric vertices as anchors to explicitly decouple and reconstruct spatial-semantic mappings for local structures. Subsequently, the AVSO performs distribution shaping and boost control over the attention, achieving active visual recall and prediction rectification.
Figure 5: The principle of the Adaptive Visual Steering Operator (AVSO).
Methods
Structure
Semantic
Overall
P
R
F1
P
R
F1
P
R
F1
Open-source MLLMs
Qwen2.5-VL-7B-Instruct
73.58
38.00
50.12
38.32
43.82
40.89
47.19
40.70
43.70
InternVL3.5-8B-Instruct
66.14
43.70
52.63
43.42
51.55
47.14
51.44
47.38
49.33
InternVL3.5-30B-Instruct
70.03
58.89
63.98
57.39
56.12
56.75
62.32
57.61
59.87
InternVL3.5-8B
68.90
54.26
60.71
61.52
59.06
60.26
64.00
56.49
60.01
Table 1: Performance comparison on PGPS9K.
Methods
Structure
Semantic
Overall
P
R
F1
P
R
F1
P
R
F1
Qwen2.5-VL-32B-Instruct
70.15
53.01
60.39
63.33
58.9
61.03
65.49
55.74
60.22
Qwen3-VL-8B-Instruct
73.16
51.04
60.13
71.15
60.15
65.19
69.89
55.15
61.65
GLM-4.6V-Flash
76.95
53.18
62.89
71.27
56.73
63.18
72.04
54.78
62.24
Qwen3.5-9B
79.72
65.21
71.74
77.98
72.45
75.11
76.70
68.48
72.36
Qwen3-VL-32B-Instruct
81.28
74.58
77.79
77.43
69.81
73.43
78.26
72.43
75.23
Table 2: Performance comparison on PGDP5K.
Methods
PGPS9K
PGDP5K
Lines
Circles
Structures
Semantics
PPR
Lines
Circles
Structures
Semantics
PPR
Qwen2.5-VL-32B-Instruct
36.57
77.51
34.75
38.42
21.43
34.64
78.64
32.60
38.42
19.36
Qwen3-VL-8B-Instruct
42.79
77.51
38.08
50.75
25.06
42.45
77.85
37.40
50.00
24.00
GLM-4.6V-Flash
43.70
80.60
39.09
49.81
25.77
46.34
80.00
39.70
51.02
25.40
Qwen3.5-9B
56.00
85.43
50.64
61.89
37.36
55.55
82.60
48.30
63.47
36.34
Qwen3-VL-32B-Instruct
64.23
92.57
62.38
58.79
43.47
64.38
87.62
58.08
58.79
40.11
Table 3: Sample Accuracy and Perfect Parsing Rate (PPR) on PGPS9K and PGDP5K.
Methods
Structure
Semantic
Overall
P
R
F1
P
R
F1
P
R
F1
Qwen3-VL-32B (Baseline)
82.91
76.16
79.39
79.63
74.67
77.07
80.37
75.48
77.85
Qwen3-VL-32B 2-stage
84.29
78.92
81.52
81.11
77.85
79.45
81.11
77.85
79.45
Uniform Intv.
86.00
80.13
82.96
79.66
77.03
78.32
81.07
78.70
79.86
w/o GCLR
86.46
81.21
83.75
81.49
78.79
80.12
82.51
80.10
81.28
w/o AVSO
86.12
81.11
83.54
80.34
78.54
79.43
81.12
79.92
80.52
Table 4: Ablation results on PGPS9K.
Figure 6: Qualitative case analysis on representative geometric diagrams.
Methods
Qwen3-VL-32B
Qwen3-VL-8B
InternVL3.5-8B
Qwen2.5-VL-7B
Qwen3-VL-4B
Direct Answer
53.73
45.10
41.18
42.16
41.18
+ Formal Language
53.73
49.61
45.29
46.08
41.37
+ Text Description
58.63
48.24
47.06
46.47
42.94
Table 5: Downstream reasoning accuracy on MathVerse.