Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On ScienceQA, Qwen2.5-VL-7B drops from 80.76% to 45.48%, and across five option-content permutations 93.54% of CoT-prefix predictions select the first slot. Condition-matched linear probes recover 78.94% from the same hidden states, while free generation restores 75.24%, showing that the answer often survives the prefix and the immediate readout fails. Vocabulary and layer diagnostics explain the mismatch: probability mass moves toward continuation tokens, while answer information remains linearly accessible in late layers. The effect recurs with varying severity across datasets and models, though not universally. These results show that CoT-prefix scoring can confound model knowledge with an evaluation-interface mismatch and should be avoided unless the requested and scored output events are aligned.
Figures & tables
Figure 1: A single suffix creates an event mismatch. Direct prompting requests and scores B; the CoT prefix requests a rationale, so immediate label scoring can return A even when free generation and a matched probe recover B.
Data
Model
N
Direct
Prefix
Δ acc
A/first
ScienceQA
Q7
4,241
80.76
45.48
−35.28
94.08
AI2D
Q7
3,088
82.48
28.47
−54.02
94.66
AI2D
Q3
3,088
79.02
46.73
−32.29
66.71
MMMU-10
Q7
286
47.90
31.82
−16.08
–
MMBench
Q7
4,377
93.31
99.11
+5.80
–
SQA*
Q3
2,017
80.07
58.11
−21.96
65.10
Table 1: Cross-dataset/model checks. Q7/Q3 denote Qwen2.5-VL-7B/3B, L7 denotes LLaVA-1.5-7B, and SQA* is the image-present subset. A/first is the CoT-prefix first-slot rate and is omitted when fixed letter labels are unavailable.
Instruction
Native Acc.
Δ acc
Native A/first
Probe Acc.
Probe A/first
None (Direct)
80.76
+0.00
52.30
84.27
35.60
Let me think step by step.
45.56
−35.20
94.08
78.94
35.96
Let’s think step by step.
49.19
−31.57
90.10
79.86
35.46
Let’s solve this step by step.
52.49
−28.27
86.49
79.96
35.32
Let’s reason step by step.
52.77
−27.99
85.76
79.89
36.95
Reasoning:
77.18
−3.58
56.17
82.60
40.30
Table 2: Trigger comparison on the ScienceQA test split. Δ acc is relative to Direct; probe columns report condition-matched accuracy and A/first rates (%).
Diagnostic
Direct
CoT-prefix
Change
Answer mass
2.15×10−8
1.51×10−9
↓14×
Reasoning mass
3.73×10−2
1.49×10−1
4.00×
Answer rank (mean)
5,251
26,562
+21,311
Reasoning top token (%)
4.3
9.9
+5.6 pp
Table 3: Full-vocabulary competition on the ScienceQA image subset ( N=2,017 ); Change compares CoT-prefix with Direct.
Applied Agent Research Center, Korea Institute of Science and Technology Information (KISTI), Republic of Korea · Department of Computer Science and Engineering, University of Seoul, Republic of Korea.