Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domain, lacking tractability due to rasterized or latent representations of images. We introduce SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks. We exploit the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process. We provide a large curated dataset of SVG-based image editing dataset, as well as the paradigm to tune open-source VLMs. Experiments on a mathematical reasoning benchmark demonstrate that SVGLM achieves strong SVG generation power as well as think-with-image intelligence. Our results highlight SVG as a suitable medium for building more robust digital domain agents, bridging the gap between text-based thinking and pixel-based images.
Figures & tables
Figure 1 : Comparison between standard multi-modal chain-of-thought and SVG-enhanced thinking. Vanilla SFT and SVGLM yield different conversation patterns and results. SVGLM offers a more robust thinking with images pipeline with better performance on reasoning tasks.
Figure 2 : Our data curation pipeline. We use Gemini-3-Pro to identify the difference between source and target images and generate SVGs. Both GPT filtering and human verification are conducted to ensure the quality of collected samples.
Perfect Reconstruction
Slight Translation
Major Translation
Structural Error
79.5%
10%
9.5%
1%
Table 1: Human evaluation results in percentage.
Model
Medium
Plane Geometry
Solid Geometry
Weighted
GPT-4o
-
18.7
20.3
19.2
V-Thinker
Code
19.0
23.8
20.5
GPT-4o + Qwen-Image-Edit
Image
19.0
20.8
19.6
LLaVa-Next-Mistral-7B
-
11.6
21.4
14.6
LLaVa-Next-Mistral-7B (SFT)
-
9.8
20.7
13.2
LLaVa-Next-Mistral-7B (SVGLM)
SVG
20.3
29.3
23.1
Table 2: Comparative evaluation of multi-modal agents on MathCanvas-Bench. Scores represent weighted accuracy. Best results within each model family are highlighted.
Figure 3 : Instruction and SVG generated by SVGLM, as well as Qwen-Image-Edit and Nano Banana on the exact same prompt
Model
Plane Geometry
Solid Geometry
Weighted
Qwen2.5-VL-7B w/o solution image
19.1
20.5
18.8
Qwen2.5-VL-7B w solution image
33.9
34.7
34.0
Table 3: Ablation study of the presence of solution images. All accuracies are calculated the same way as Table 2 .
Model
Zero-shot
SFT
SVGLM
LLaVa-Next-Mistral-7B
47.4
61.2
69.4
Qwen2.5-VL-7B
74.4
81.9
83.8
InternVL3-8B
81.6
82.3
84.8
Table 4: Evaluation results of grounded reasoning with POPE benchmark
Model
SFT
SVGLM
w/o SVG render
w/o SVG gen. & render
LLaVa-Next
11.6
20.0
19.5
10.7
Qwen2.5-VL
20.3
26.8
23.8
20.2
InternVL3
21.7
26.2
23.8
21.4
Table 5: Ablation study verifying all framework components. All figures represent the weighted score on MathCanvas-Bench. w/o SVG render forces the model to output the answer after writing the SVG code, without the rendered image as feedback. w/o SVG gen. & render represents the baseline without any SVG assistance.
Model
No Reasoning
Concise
Complete
LLaVa-Next
18.0
20.0
19.0
Qwen2.5-VL
26.8
23.9
23.4
InternVL3
26.2
24.8
22.4
Table 6: Ablation study of reasoning density. All figures are the weighted score of MathCanvas-Bench.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Template used for SVG generation task. The text highlighted in cyan indicates variables to be replaced.
Figure 5 : Template used for SVG quality assessment.
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink is a think-with-SVG pipeline, where the model learns to generate, render, and interpret scalable vector graphics (SVG) code as an intermediate visual representation for multi-turn reasoning. By creating structured vector sketches, the model can externalize spatial hypotheses, inspect them through deterministic rendering, and reason within a constrained geometric space, effectively mimicking the human process of mental imagery. We instantiate this paradigm through a two-stage training framework, combining Supervised Fine-Tuning (SFT) for SVG syntactic alignment with multi-turn Reinforcement Learning (RL) to encourage iterative inspection, revision, and refinement of intermediate visual hypotheses. Extensive evaluations demonstrate that MentalThink achieves superior performance on spatial understanding and reasoning benchmarks (e.g., 55.1% on VSIBench, 76.0% on MindCube), showing that executable vector graphics provide a verifiable visual workspace for dynamic perspective taking, visual reflection, and compositional scene construction.
Kangheng Lin, Jisheng Yin, Dingming Li +11
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications · University of Chinese Academy of Sciences · StepFun
Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introduce errors and increasingly interfere with textual reasoning as reasoning progresses. We propose Visual-to-Text Chain-of-Thought Distillation (V2T), a framework that enables LVLMs to internalize visual reasoning without generating intermediate visual representations at inference time. V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning. When reasoning images can be mapped to the original image, V2T can additionally distill the teacher's attention to corresponding regions, while ground-truth bounding boxes can further guide a subsequent reinforcement learning stage. Experiments across multiple multimodal reasoning benchmarks show that V2T consistently outperforms the teacher and existing baselines, improving average accuracy by 14.3% on a held-out set and 2.7% on the broader visual evaluation suite. Moreover, lightweight SFT and substantially reduced RL make V2T up to 42x faster to train than state-of-the-art baselines.
Wenhan Yang, Nilay Naharas, Ali Payani +1
University of California, Los Angeles · Cisco Research
When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models (VLMs) such as Gemini-3-Pro and GPT-5 only respond with text, which can be difficult for users to verify. We present SketchVLM, a training-free, model-agnostic framework that enables VLMs to produce non-destructive, editable SVG overlays on the input image to visually explain their answers. Across seven benchmarks spanning visual reasoning (maze navigation, ball-drop trajectory prediction, and object counting) and drawing (part labeling, connecting-the-dots, and drawing shapes around objects), SketchVLM improves visual reasoning task accuracy by up to +28.5 percentage points and annotation quality by up to 1.48x relative to image-editing and fine-tuned sketching baselines, while also producing annotations that are more faithful to the model's stated answer. We find that single-turn generation already achieves strong accuracy and annotation quality, and multi-turn generation opens up further opportunities for human-AI collaboration. An interactive demo and code are at https://sketchvlm.github.io/.