cs.CVOct 7, 2026

From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

Authors: Zijian Chen, Zhengyu Chen, Bohan Liang, Lirong Deng, Yushuo Zheng, Yanwei Jiang, Qi Jia, Kaiwei Zhang, +2 more

Organizations: Institute of Image Communication and Information Processing, Shanghai Jiao Tong University, Shanghai, 200240, China. · Shanghai Artificial Intelligence Laboratory, Shanghai, 200030, China. · Macao Polytechnic University, Macao, 999078, China.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.

Figures & tables

Explore similar work

CardsList
  1. From Charts to Code: A Hierarchical Benchmark for Multimodal Models

    Oct 20, 2025Jiahao Tang, Henry Hengyuan Zhao, Lijian Wu +8Image-To-Code GenerationChart

  2. Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

    Jun 14, 2026Xuanle Zhao, Qiushi Sun, Jingyu Xiao +16Survey