cs.CVSep 28, 2026

Distilling Visual Reasoning into Text Space

Authors: Wenhan Yang, Nilay Naharas, Ali Payani, Baharan Mirzasoleiman

Organizations: University of California, Los Angeles · Cisco Research

Abstract

Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introduce errors and increasingly interfere with textual reasoning as reasoning progresses. We propose Visual-to-Text Chain-of-Thought Distillation (V2T), a framework that enables LVLMs to internalize visual reasoning without generating intermediate visual representations at inference time. V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning. When reasoning images can be mapped to the original image, V2T can additionally distill the teacher's attention to corresponding regions, while ground-truth bounding boxes can further guide a subsequent reinforcement learning stage. Experiments across multiple multimodal reasoning benchmarks show that V2T consistently outperforms the teacher and existing baselines, improving average accuracy by 14.3% on a held-out set and 2.7% on the broader visual evaluation suite. Moreover, lightweight SFT and substantially reduced RL make V2T up to 42x faster to train than state-of-the-art baselines.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding

    Jun 7, 2026Lianyu Hu, Xiaoyu Ma, Zeqin Liao +1Chain-of-Thought Reasoning

  2. Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs

    Apr 8, 2026Mengdan Zhu, Senhao Cheng, Liang ZhaoRecent Vision-Language ModelsVisual Reasoning

  3. UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs

    May 12, 2026Houcheng Jiang, Jiajun Fu, Junfeng Fang +4Latent Visual ReasoningMultimodal Large Language Models