Seeing and Solving Are Not Enough for Vision-Language Models
Authors: Ziheng Wang, Mingxuan Xie, Yilin Liu, Dayan Wu, Yang Li, Pengwen Dai
Organizations: Sun Yat-sen University · Zhejiang University · The Hong Kong University of Science and Technology · Institute of Information Engineering · Hunan University
Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap.
Figures & tables
Figure 1: An example of a composition failure. For the same question, the model succeeds at both Extract and Solve, yet fails to answer directly. Panel (a) shows the task-state components schematically. In the actual evaluation, Extract and Direct each receive the same full task question in a single call, without the illustrative scene-content lists (Appendix G ).
Figure 2: State Realization Tuning with LoRA. Only the SRT adapter, implemented as LoRA updates in the language-model layers, is trained. The task state and final answer are two fields of one autoregressive response, with the state generated first.
Accuracy
Failures
Error breakdown
Dataset
Model
Extract
Solve
Direct
CFS
CFR
ChartQA
Qwen
81.1
86.8
63.5
40.2
20.5
Molmo
75.5
94.3
58.0
65.5
38.2
InternVL
90.3
97.6
66.6
75.6
28.7
MiniCPM
77.2
98.1
61.6
56.5
28.8
CLEVR
Qwen
56.3
94.5
37.5
46.1
54.2
Table 1: Accuracy and composition-failure prevalence. CFS is the share of Direct errors that are composition failures. CFR is the Direct failure rate on questions where both Extract and Solve succeed. The bars partition each model’s Direct errors. All values are percentages.
Variants
Dataset
Model
Untuned
Standard SFT
Format Control
Answer-to-State
SRT
ChartQA
Qwen
63.5
85.1±0.8
85.8±0.7
85.1±0.6
91.7±0.4
Molmo
58.0
77.3±0.6
75.7±2.4
79.4±1.0
91.5±0.6
InternVL
66.6
87.7±0.8
86.2±0.8
87.4±1.2
94.0±0.4
MiniCPM
61.6
72.8±1.3
72.3±1.7
73.0±1.0
86.0±0.4
CLEVR
Qwen
37.5
93.7
92.9
93.5
95.4
Table 2: Answer accuracy under various conditions. Untuned is the model before fine-tuning. ChartQA results are the mean ± standard deviation over three seeds; CLEVR results use one seed. All values are percentages.
Figure 3: A qualitative example of SRT on CLEVR. Standard SFT produces an incorrect answer, whereas SRT first realizes the correct task state and then derives the correct answer.
State realization
CF Repair
Preservation
Dataset
Model
State Acc.
S–A Cons.
Standard SFT
SRT
ChartQA
Qwen
89.1±0.2
99.2±0.5
73.0±4.5
94.3±1.1
98.8±0.1
Molmo
88.1±0.8
99.3±0.1
71.4±0.2
98.1±1.4
98.5±0.6
InternVL
91.7±0.6
99.8±0.2
82.4±1.7
97.5±0.4
99.2±0.3
MiniCPM
81.2±1.0
99.8±0.1
58.3±3.0
92.5±2.4
97.1±0.1
CLEVR
Qwen
93.9
100.0
95.3
96.8
97.0
Table 3: State realization, failure repair, and preservation. State metrics and Preservation refer to SRT. CF Repair compares Standard SFT and SRT on fixed pre-tuning composition-failure sets. Values are percentages, with three-seed standard deviations on ChartQA.
Model
Untuned
Standard SFT
Extract → Solve
SRT
Qwen
63.5
85.1±0.8
75.1
91.7±0.4
Molmo
58.0
77.3±0.6
79.7
91.5±0.6
InternVL
66.6
87.7±0.8
90.7
94.0±0.4
MiniCPM
61.6
72.8±1.3
82.1
86.0±0.4
Table 4: Inference-time composition versus supervised fine-tuning on ChartQA. Extract → Solve passes the predicted state to a separate image-free Solve call. Fine-tuning results are three-seed mean ± standard deviation. All values are percentages.
Task structure
Qwen
Molmo
Multi-step arithmetic
96.2 (25/26)
86.0 (49/57)
Conditional lookup
50.8 (33/65)
50.0 (2/4)
Relational count ranking
80.0 (28/35)
N/A (0/0)
Color-set operations
45.0 (68/151)
56.8 (25/44)
Overall CFR
55.6 (154/277)
72.4 (76/105)
Table 5: Composition failures with richer task states. Untuned CFR (%) is shown with CF /J counts. N/A means no question passes both Extract and Solve ( J=0 ).
Qwen
Molmo
Task structure
N
Standard SFT
SRT
Standard SFT
SRT
Multi-step arithmetic
291
32.3
67.7
22.7
68.7
Conditional lookup
215
71.6
85.1
67.4
84.2
Relational count ranking
400
87.8
96.5
38.5
80.0
Color-set operations
400
79.5
98.5
73.5
91.0
Average accuracy
1,306
70.2
88.8
50.5
81.5
Table 6: SRT on additional task structures. Final-answer accuracy (%) of the Standard SFT and SRT adapters on each structure’s N test questions.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Model
N
J
CF
CFS
CFR
ChartQA
Qwen
1,077
770
158
40.2
20.5
ChartQA
Molmo
1,077
774
296
65.5
38.2
ChartQA
InternVL
1,077
948
272
75.6
28.7
ChartQA
MiniCPM
1,077
813
234
56.5
28.8
CLEVR
Qwen
33,723
17,916
9,708
46.1
54.2
CLEVR
Molmo
33,723
7,014
4,044
17.7
57.7
Appendix
Table 7: Exact counts underlying the composition-failure diagnosis. J is the number of questions on which both Extract and Solve succeed; CF is the subset for which Direct fails.
Setting
Model
Direct errors
E=0,S=0
E=0,S=1
E=1,S=0
E=1,S=1
ChartQA
Qwen
393
38 (9.7)
106 (27.0)
91 (23.2)
158 (40.2)
ChartQA
Molmo
452
13 (2.9)
133 (29.4)
10 (2.2)
296 (65.5)
ChartQA
InternVL
360
2 (0.6)
74 (20.6)
12 (3.3)
272 (75.6)
ChartQA
MiniCPM
414
3 (0.7)
165 (39.9)
12 (2.9)
234 (56.5)
CLEVR
Qwen
21,076
783 (3.7)
9,550 (45.3)
1,035 (4.9)
9,708 (46.1)
Appendix
Table 8: Direct errors broken down by Extract and Solve outcomes. Counts, with the percentage of Direct errors in parentheses. The last column is the composition-failure set C ; the first two together form E and the third is S in Equation 4 .
Model
Operation
CF count
CFS
CFR
Qwen
Comparison
53
29.0
26.4
Difference
55
52.4
19.9
Sum
50
47.6
17.1
Molmo
Comparison
60
75.0
24.8
Difference
107
59.8
43.1
Sum
129
66.8
45.4
Appendix
Table 9: Composition failures by operation on ChartQA. CFS and CFR are percentages.
Setting
Model
Direct (strict)
Direct (content)
CFS (content)
CFR (content)
ChartQA
Molmo
58.0
61.6
63.5
34.0
ChartQA
MiniCPM
61.6
61.6
56.5
28.8
CLEVR
Qwen
37.5
38.6
45.5
52.6
Appendix
Table 10: Sensitivity to content-normalized Direct scoring. Extract and Solve outcomes are held fixed; only final-answer presentation is normalized.
Model
Canonical CFR
Paraphrase A
Paraphrase B
Molmo
34.0
34.5
30.2
MiniCPM
28.8
28.8
27.6
Appendix
Table 11: Composition failures persist across prompt paraphrases. Values are CFR (%); all three wording sets use content-normalized scoring.
Setting
Model
Reversed
J
CF
CFR
ChartQA
Qwen
10
770 → 780
158 → 158
20.5 → 20.3
ChartQA
Molmo
31
774 → 805
296 → 303
38.2 → 37.6
ChartQA
InternVL
3
948 → 951
272 → 272
28.7 → 28.6
ChartQA
MiniCPM
10
813 → 823
234 → 236
28.8 → 28.7
CLEVR
Qwen
212
17,916 → 18,128
9,708 → 9,793
54.2 → 54.0
Appendix
Table 12: Sensitivity of the diagnosis to accepting reversed pairs on sum and absolute-difference questions. “Reversed” counts Extract outputs that fail strict scoring only because the two values are swapped. Strict → relaxed values are shown for J , CF, and CFR (%).
Model
Solve (text-only)
Image + ground-truth task state
Image + counterfactual state
Qwen
86.8
85.4
45.7
Molmo
94.3
92.1
38.4
Appendix
Table 13: Solve accuracy on ChartQA with and without the image. A task state is supplied in text in each condition. The counterfactual condition replaces the ground-truth state with an answer-changing incorrect state. Invalid-output rates for Molmo are 3.0% (image + ground-truth) and 7.9% (image + counterfactual). Values are percentages.
Setting
Model
Standard SFT
Format Control
Answer-to-State
SRT
Δ Format
Δ Ans.-to-State
ChartQA
Qwen
85.1
85.8
85.1
91.7
+5.9
+6.7
ChartQA
Molmo
77.3
75.7
79.7
91.6
+15.9
+11.9
ChartQA
InternVL
87.7
86.2
87.4
94.0
+7.8
+6.6
ChartQA
MiniCPM
72.8
72.3
73.0
86.0
+13.7
+13.0
CLEVR
Qwen
93.7
92.9
93.5
95.4
+2.5
+1.9
Appendix
Table 14: Final-answer accuracy under content-only scoring. The auxiliary field is ignored. ChartQA values are three-seed means. The last two columns show SRT’s gains over the two controls in percentage points, computed from unrounded means.
Setting
Model
Seed
Wrong state, answer correct
Correct state, answer wrong
ChartQA
Qwen
13
36 / 120
10 / 957
42
36 / 115
7 / 962
87
33 / 117
3 / 960
Molmo
13
38 / 125
1 / 951
42
37 / 122
3 / 955
87
40 / 132
1 / 939
Appendix
Table 15: Generated states and final answers under SRT. Each fraction counts the indicated answer outcome among questions with the indicated state outcome. ChartQA Molmo outputs with an undefined state (1, 0, and 6 per seed) are excluded. Additional task-structure results use the Qwen adapter from Table 6 .
Figure 4: State correctness and state–answer consistency are distinct. (a) On ChartQA, Qwen gives an incorrect answer that is consistent with its incorrect generated state. (b) Molmo extracts the three values correctly but computes the wrong answer. The latter example comes from the additional multi-step arithmetic experiment. Questions and outputs are verbatim.
Model
Direct
Zero-shot State-first
Invalid
SRT
Qwen
63.5
59.1
0.1
91.7
Molmo
58.0
53.4
5.7
91.5
InternVL
66.6
11.6
78.8
94.0
MiniCPM
61.6
53.9
13.6
86.0
Appendix
Table 16: Zero-shot state-first prompting on ChartQA. Direct and zero-shot answers use the same canonical final-answer parser. Invalid zero-shot outputs count as incorrect. SRT gives the corresponding three-seed fine-tuning mean. All values are percentages.
Model
Family
N
J
CF
CFS
CFR
95% CI
Qwen
Multi-step
291
26
25
9.2
96.2
[80.4, 99.9]
Conditional lookup
215
65
33
26.0
50.8
[38.1, 63.4]
Relational count ranking
400
35
28
10.9
80.0
[63.1, 91.6]
Color-set
400
151
68
31.5
45.0
[36.9, 53.3]
Molmo
Multi-step
291
57
49
17.4
86.0
[74.2, 93.7]
Conditional lookup
215
4
2
1.1
50.0
[6.8, 93.2]
Appendix
Table 17: Untuned diagnosis on the additional task structures. J counts questions where Extract and Solve both succeed; CF counts composition failures. The last column gives the exact two-sided 95% Clopper–Pearson interval for CFR. N/A means that CFR and its interval cannot be computed because J=0 . These intervals describe uncertainty over the evaluated questions, not variation across training seeds.
Model
Family
State Acc.
S–A Cons.
Qwen
Multi-step
84.2
81.8
Conditional lookup
79.5
96.2
Relational count ranking
78.5
100.0
Color-set
97.5
99.3
Molmo
Multi-step
80.8
83.7
Conditional lookup
78.6
99.1
Appendix
Table 18: State realization on the additional task structures. Results use the multi-task SRT adapters from Table 6 . State Acc. is exact task-state accuracy over all questions. S–A Cons. measures whether the answer follows from the generated state, among outputs where the state defines an answer and the final answer can be parsed. Values are percentages.
Figure 5: A qualitative example of SRT on ChartQA. For Qwen, Standard SFT produces an incorrect answer, whereas SRT first realizes the correct task state and then computes the sum correctly.
Figure 6: Relational count ranking and color-set difference. SRT lists the three relation-conditioned counts before choosing the middle group, and both color sets before taking their difference.
Figure 7: Three-value arithmetic and conditional lookup. SRT states three numerical values or a lookup table before answering.
Figure 8: State realization across chart appearances and operations. Qwen compares two values in a line chart and computes an absolute difference in a stacked-bar chart. Molmo states three values before a multi-step calculation on a different line chart. Panel (c) belongs to the additional task structures, not the pair-state ChartQA comparison. All outputs are verbatim.
Figure 9: Composition within the CLEVR count queries. The first query combines material and a union of object groups, while the second uses a spatial relation. SRT gives both counts and their correct sum, whereas Standard SFT answers incorrectly. Question and outputs are verbatim from Qwen.
Condition
Image
Full question
State supplied
Output
Extract
Yes
Yes
No
Task state
Solve
No
Yes
Yes
Final answer
Direct
Yes
Yes
No
Final answer
Standard SFT
Yes
Yes
No
Final answer
Format Control
Yes
Yes
No
Fixed placeholder, then answer
Answer-to-State
Yes
Yes
No
Answer, then task state
Appendix
Table 19: Information provided to each condition. The state column indicates whether the ground-truth task state is supplied as input. For fine-tuning, the output column gives the assistant supervision target; at test time the model generates that output itself. All conditions receive the full task question.
Method
Assistant target
Standard SFT
FINAL_ANSWER: y
Format Control
CHECK_STATE: X FINAL_ANSWER: y
Answer-to-State
FINAL_ANSWER: y TASK_STATE: a , b
SRT
TASK_STATE: a , b FINAL_ANSWER: y
Appendix
Table 20: Training targets for the pair-state tasks. Here a,b are the ground-truth values or counts, y is the answer, and X is the fixed, task-independent placeholder. Field names are literal, including the underscore. Each displayed line is one output line.