cs.CVSep 24, 2026

Multimodal Thinking with Renderable Programs

Authors: Sunli Chen, Ding Zhong, Ziqiao Ma, Jiaxin Liu, Zeyuan Yang, Hao Zhang, Lie Lu, Joyce Chai, +1 more

Organizations: University of Massachusetts Amherst · University of Michigan · University of Illinois Urbana-Champaign · Dolby Laboratories

Abstract

Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domain, lacking tractability due to rasterized or latent representations of images. We introduce SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks. We exploit the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process. We provide a large curated dataset of SVG-based image editing dataset, as well as the paradigm to tune open-source VLMs. Experiments on a mathematical reasoning benchmark demonstrate that SVGLM achieves strong SVG generation power as well as think-with-image intelligence. Our results highlight SVG as a suitable medium for building more robust digital domain agents, bridging the gap between text-based thinking and pixel-based images.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MentalThink: Shaping Thoughts in Mental SVG World

    Jul 3, 2026Kangheng Lin, Jisheng Yin, Dingming Li +11ThinkImagination

  2. Distilling Visual Reasoning into Text Space

    Sep 28, 2026Wenhan Yang, Nilay Naharas, Ali Payani +1Recent Vision-Language ModelsMultimodal Reasoning Benchmarks

  3. SketchVLM: Vision language models can annotate images to explain thoughts and guide users

    Apr 23, 2026Brandon Collins, Logan Bolton, Hung Huy Nguyen +3Visual ReasoningSketches