cs.CLSep 24, 2026

Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

Authors: Zeyan Li, Siyuan Qiu, Jianfeng Xu

Organizations: Shanghai Jiao Tong University

Abstract

Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On ScienceQA, Qwen2.5-VL-7B drops from 80.76% to 45.48%, and across five option-content permutations 93.54% of CoT-prefix predictions select the first slot. Condition-matched linear probes recover 78.94% from the same hidden states, while free generation restores 75.24%, showing that the answer often survives the prefix and the immediate readout fails. Vocabulary and layer diagnostics explain the mismatch: probability mass moves toward continuation tokens, while answer information remains linearly accessible in late layers. The effect recurs with varying severity across datasets and models, though not universally. These results show that CoT-prefix scoring can confound model knowledge with an evaluation-interface mismatch and should be avoided unless the requested and scored output events are aligned.

Figures & tables

Explore similar work

CardsList
  1. Reliable Chain-of-Thought via Prefix Consistency

    May 8, 2026Naoto Iwase, Yuki Ichihara, Mohammad Atif Quamar +1Self-ConsistencyChain-of-Thought Reasoning

  2. CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

    Jun 16, 2026Sneha Rao, Shaina Raza, Dhanesh RamachandramRecent Vision-Language Models

  3. TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

    May 28, 2026Yundong Kim, Heyoung YangLLM Reasoning StrategiesLarge Language Model Evaluation