Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introduce errors and increasingly interfere with textual reasoning as reasoning progresses. We propose Visual-to-Text Chain-of-Thought Distillation (V2T), a framework that enables LVLMs to internalize visual reasoning without generating intermediate visual representations at inference time. V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning. When reasoning images can be mapped to the original image, V2T can additionally distill the teacher's attention to corresponding regions, while ground-truth bounding boxes can further guide a subsequent reinforcement learning stage. Experiments across multiple multimodal reasoning benchmarks show that V2T consistently outperforms the teacher and existing baselines, improving average accuracy by 14.3% on a held-out set and 2.7% on the broader visual evaluation suite. Moreover, lightweight SFT and substantially reduced RL make V2T up to 42x faster to train than state-of-the-art baselines.
Figures & tables
Figure 2: Interleaved visual inputs increasingly interfere with textual reasoning. (a–c) The loss gap between reasoning with and without interleaved images widens as reasoning progresses and further increases with irrelevant visual content. (d) With longer reasoning prefixes, text-only continuation matches or outperforms continuation with interleaved images.
Figure 3: V2T Training. V2T first trains a teacher with interleaved images to ground its CoT in visual cues. The student model, which does not have access to interleaved images, is trained using KL divergence to match the teacher’s textual CoT and logits, along with attention distillation to transfer the teacher’s visual grounding, and cross-entropy loss to match the ground-truth textual CoT.
Figure 4: V2T inference does not require intermediate images. At inference, the base model, teacher (Interleaved CoT SFT), and V2T student (without attention distillation or GRPO) receive the question and original problem image and generate textual responses. The base model and teacher misidentify the object as a walking stick, while the student correctly identifies a baseball bat. Phrases such as “zoomed-in view” are inherited from the teacher’s interleaved CoT format; no intermediate image is generated or provided, and V2T reasons only over the original image.
Method
Graph Algorithm
RPM
Tetris
Visual Search
Avg
LVLMs with text-only visual CoT
R1-VL-7B
44.5
10.6
55.0
60.5
42.6
VLAA-Thinker-7B
58.0
54.8
68.0
63.0
60.9
VL-Rethinker-7B
45.0
55.8
69.0
64.5
58.6
LVLMs with interleaved latent visual CoT
LVR-7B
17.5
23.1
25.5
66.5
33.2
Table 1: Zebra-CoT held-out set. Best and second-best results across all methods are highlighted in bold and underline , respectively. As Bagel-Zebra is trained on the full Zebra-CoT dataset, we exclude it from the evaluation. V2T achieves the best overall performance of 75.5% , outperforming the interleaved teacher model by 14.3% .
Figure 5: Qualitative comparison of attention maps among Text-Only SFT, Interleaved SFT, and V2T on the Visual Search split of Zebra-CoT. V2T attends more strongly to task-relevant visual regions, suggesting improved visual grounding compared to text-only and Interleaved SFT (teacher) models.
Category
Text-Only SFT
Interleaved SFT
V2T (Ours)
Perception & Grounding
366
382
364
Spatial & Global Context
241
230
193
Reasoning & Knowledge
941
916
872
Task & Answer Consistency
199
209
179
Total
1747
1737
1608
Table 5: Error distribution. V2T has the largest relative reduction in Spatial & Global Context errors.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
LoRA rank
32
LoRA alpha
64
LoRA target modules
All
Learning rate
1.0×10−4
Per-device batch size
4
Gradient accumulation steps
4
Appendix
Table 6: Hyperparameters for SFT.
Hyperparameter
Value
Initialization
V2T student
Training data
TreeVGR-RL-37K
Fine-tuning
Full-parameter
Learning rate
1.0×10−6
Learning-rate scheduler
Constant
KL coefficient (low-variance KL)
0.01
Appendix
Table 7: Hyperparameters for GRPO.
Subset
Percentage
Visual logic and strategic games – ARC-AGI-New
0.056%
Scientific reasoning – Geometry
1.045%
Scientific reasoning – Competitive Programming
1.040%
Visual logic and strategic games – RPM
2.266%
Visual logic and strategic games – ARC-AGI
0.056%
3D visual reasoning – Embodied CoT
0.172%
Appendix
Table 8: Composition of the filtered Zebra-CoT training set.
Figure 6: Hyperparameter ablations for V2T and V2T+GRPO. Left: average accuracy of V2T when varying the logit distillation weight λ and the attention supervision weight β . Right: average accuracy of V2T+GRPO when varying the visual grounding reward weight η .
Model
MathVerse
MathVista
MM-Math
DynaMath
Qwen2.5-VL-7B
43.9
69.1
41.1
56.5
VLAA-Thinker-7B
45.8
71.4
43.1
57.3
Appendix
Table 9: Performance of Qwen2.5-VL-7B and VLAA-Thinker-7B on math-focused reasoning benchmarks.
Model
Judge
HR-Bench 4K
HR-Bench 8K
V*Bench
Qwen2.5-VL-7B
Qwen3-235B
68.9
64.1
75.9
Qwen2.5-VL-7B
GPT-4o
68.9
64.0
75.9
V2T + GRPO
Qwen3-235B
73.5
71.3
82.2
V2T + GRPO
GPT-4o
73.1
71.2
81.7
Appendix
Table 10: Judge-consistency comparison on selected benchmarks. GPT-4o is used only for this cross-check; Qwen3-235B-A22B-Instruct-2507-FP8 remains the judge used in our main evaluation.
Figure 7: Qualitative comparison on a traffic-light recognition example. The top panel shows the original question image. The base model predicts Green . The interleaved SFT model also predicts Green after performing step-wise reasoning over the scene. In contrast, V2T focuses on the distant traffic lights and correctly predicts Red , matching the ground-truth answer. Phrases such as “zoomed-in view” are inherited from the teacher’s interleaved CoT format; no intermediate image is generated or provided, and V2T reasons only over the original image.
Figure 8: Qualitative comparison on a counting example. The question image contains two visible white flowers in the central arrangement. The base model and the interleaved SFT model both under-count the flowers and predict only one white flower. In contrast, V2T focuses on the central floral arrangement, identifies two distinct white flowers, and correctly predicts the ground-truth answer. Phrases such as “zoomed-in view” are inherited from the teacher’s interleaved CoT format; no intermediate image is generated or provided, and V2T reasons only over the original image.
Figure 9: Generating intermediate visual states and interleaving grounded reasoning for chemistry.
Figure 10: Generating intermediate visual states and interleaving grounded reasoning for chess.
Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (MLLMs), existing CoT approaches suffer from a fundamental limitation: they perform reasoning entirely in text without accessing visual features during the reasoning process. After initial visual encoding, image information becomes inaccessible, forcing models to reason based solely on whatever was captured in the initial description, which forms a `vision-blind reasoning' paradigm that limits fine-grained visual extraction, error verification, and adaptive attention. We propose Text-Visual Interleaved Chain-of-Thought (TVI-CoT), a framework that enables explicit interleaving of textual reasoning and visual feature access through learnable control tokens <THINK>, <LOOK> and <ANSWER>. These tokens allow dynamic switching between reasoning and visual grounding, attending to relevant image regions conditioned on the evolving reasoning state. Experiments on eight benchmarks demonstrate state-of-the-art results among MLLM-based CoT methods and notable performance boost compared to the baseline: +6.1% on MMMU, +3.8% on MathVerse, +3.4% on MathVista, and +3.4% on ScienceQA. Code is available at https://github.com/hulianyuyy/TVI-CoT.
Lianyu Hu, Xiaoyu Ma, Zeqin Liao +1
College of Computing and Data Science of Nanyang Technological University, Singapore.
Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that are insufficient to extract semantics in multi-step reasoning. We propose "Decompose, Look, and Reason" (DLR), a reinforced latent reasoning framework that dynamically decomposes queries into textual premises, extracts premise-conditioned continuous visual latents, and deduces answers through grounded rationales. We introduce a three-stage training pipeline and propose a novel Spherical Gaussian Latent Policy, to enable effective exploration in the latent space. Extensive experiments on vision-centric benchmarks show that DLR consistently outperforms strong baselines, including text-only, interleaved multimodal CoT, and latent reasoning methods, while providing superior stepwise interpretability.
Mengdan Zhu, Senhao Cheng, Liang Zhao
Emory University · University of Michigan, Ann Arbor
Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on explicit textual chain-of-thought interleaved with visual latent tokens. This interleaved design limits efficiency and keeps reasoning fragmented across separate text and vision channels. We propose UniVLR, a unified visual latent reasoning framework that treats textual reasoning and auxiliary visual evidence as a shared visual workspace. Instead of preserving text CoT as an independent inference-time path, UniVLR renders reasoning traces together with auxiliary images and learns to compress this unified representation into compact visual latent tokens. At inference time, the model reasons only through visual latents and directly decodes the final answer, avoiding both external tool calls and verbose text reasoning. Experiments on real-world perception and visual reasoning tasks show that UniVLR outperforms prior visual latent reasoning methods while using substantially fewer generated reasoning tokens, suggesting a more unified and efficient paradigm for visual thinking in MLLMs.
Houcheng Jiang, Jiajun Fu, Junfeng Fang +4
University of Science and Technology of China · Zhongguancun Academy · National University of Singapore +1