Reinforcement learning is increasingly used to post-train vision-language models for image-to-code generation, such as generating SVG code from a reference image, by optimizing rewards computed from the final rendered output. However, relying on a single terminal reward provides sparse feedback that is poorly aligned with the contribution of individual tokens. A generated program may contain operations that accurately reproduce some parts of the target image alongside others that introduce errors, yet all tokens are trained from the same final outcome. We observe that many intermediate code prefixes are not only executable, but already produce meaningful partial renders that reflect progress toward the target. This property provides a natural source of denser supervision during generation. Based on this observation, we introduce IR4RL, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence. We evaluate our approach on Image-to-SVG and Image-to-TikZ generation. Across both tasks, our method improves over supervised fine-tuning and standard GRPO, yielding new state-of-the-art open-source models. This shows that intermediate rendering provides a simple and effective source of process supervision for RL post-training of image-to-code models.
Figures & tables
Figure 1: We introduce a method that leverages I ntermediate R enders for R einforcement L earning ( IR4RL ) post-training of image-to-code VLMs as a supervision signal, allowing us to improve over outcome-only RL and achieve state-of-the-art results in Image-to-SVG and Image-to-TikZ tasks.
Figure 2: Process Reward from Intermediate Renders. Rollouts 1 and 2 reach similar final quality through different generation trajectories while Rollout 3 makes early progress that later operations reverse. Outcome-only advantage rewards all tokens within a rollout equally without distinguishing helpful from harmful decisions. Our method complements it with localized feedback leveraging changes in visual score between consecutive renders (the arrows Δj ).
Figure 3: Qualitative evaluation of Image-to-SVG. We compare LIVE, Gemini-3-Flash, task-specific SVG models, and post-training baselines. Our method better preserves structure, color, geometry, and fine details, while competing methods more often omit or distort visual elements.
MMSVGBench-Illustrations
MMSVGBench-Icons
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
Optimization-based
DiffVG
92.03
9.23
0.35
94.10
93.52
4.85
79.6k
90.97
9.24
0.44
93.76
95.29
4.89
79.5k
LIVE
94.55
10.02
0.72
95.48
93.68
4.99
8.4k
94.24
9.18
0.86
95.19
95.60
4.91
8.4k
General-purpose (M)LLMs
Qwen3-VL-235B
92.81
28.32
5.30
87.89
89.82
4.82
5.0k
92.23
29.90
7.78
84.77
91.61
4.77
5.9k
Table 1: Quantitative evaluation on MMSVGBench. We compare optimization-based methods, general-purpose VLMs, SVG models, and post-training baselines on the Illustrations and Icons splits. Our method surpasses baselines, substantially improves the SFT base model, and performs best when process and outcome supervision are combined. Bold and underline denote the best and second-best results among learned methods, excluding optimization-based methods.
Gemini 3 Flash
InternSVG-8B
OmniSVG-4B (SFT)
OmniSVG-4B + RAFT
OmniSVG-4B + Outcome RL
Users ↑
60.6%
76.3%
92.7%
74.9%
74.0%
VLM ↑
54.0%
74.0%
94.0%
78.0%
76.0%
Table 2: User Study and VLM evaluation. 2AFC win rate (%) of our method vs. each baseline demonstrates our method is preferred by humans and a VLM (VLM-human agreement: 88%).
Figure 5: Qualitative reward composition ablation. Process-only training recovers much of the gain over outcome-only, while combining both rewards yields the most faithful reconstructions.
DreamSim ↑
SigLIP ↑
CLIP ↑
LPIPS ↓
KID ↓
C-BLEU ↑
TED ↓
Tokens ↓
General-purpose (M)LLMs
Qwen3-VL-235B
80.3
91.5
89.0
40.4
0.74
3.4
54.1
3.9k
Gemini 3 Flash
91.0
96.1
94.3
27.2
-0.05
6.6
51.7
2.4k
Sonnet 5
85.4
94.0
92.2
35.9
0.06
3.7
52.8
0.7k
GPT-5.2
84.9
93.8
91.4
35.7
0.41
4.4
54.0
1.5k
TikZ VLMs
Table 3: Quantitative evaluation on DaTikZ-v3. Our method improves DeTikZify-v2 and post-training baselines across the visual reconstruction metrics while producing substantially shorter programs. Bold and underline mark best and second-best results among task-specific TikZ models.
Figure 6: Ablations of process supervision. (a) Increasing the process contribution improves performance up to α=10 . (b) Reward propagation performs best at λ=0.9 . (c) More frequent intermediate renders consistently improve performance. Stars mark our default settings.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1: Best-of- K sampling on SVG generation. For each image, we sample K candidate SVGs and report the mean reward of the best candidate. Across multiple datasets and temperature values, our method performs best across all sampling budgets, with the largest gains at small K .
Figure A.2: Prefix closure for (a) Image-to-SVG, (b) Image-to-TikZ. A model generates a program code y that is separated into granular segments y1:bj , well-formed by a closure operator C into Pj=C(y1:bj) and rendered. This process depends on the task and a tokenizer. Specifically, OmniSVG does not generate the <svg> tag, while in DeTikZify the opening code is part of the sequence.
MMSVGBench-Illustrations
MMSVGBench-Icons
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
700 samples
97.48
9.79
1.17
94.22
96.05
5.09
2.5k
98.26
8.67
1.23
93.88
98.01
4.98
2.1k
10k samples
97.26
10.11
1.27
94.11
95.35
5.08
2.2k
98.49
8.42
1.20
94.12
98.35
5.02
1.9k
Appendix
Table A.1: Effect of the number of training samples. We observe that increasing the size of the trainset provides only marginal gains on Icons subset at the cost of mild degradation on Illustrations, and increases the overall generation length. We thus opt for a smaller trainset in all experiments.
MMSVGBench-Illustrations
MMSVGBench-Icons
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
OmniSVG-4B
85.48
22.26
5.11
89.70
82.03
4.55
11.3k
89.21
19.82
5.80
87.96
88.42
4.65
8.4k
OmniSVG-8B
88.81
21.15
4.98
88.85
85.57
4.68
9.1k
92.04
18.09
4.92
89.28
91.89
4.79
6.0k
OmniSVG-4B + Ours
97.48
9.79
1.17
94.22
96.05
5.09
2.5k
98.26
8.67
1.23
93.88
98.01
4.98
2.1k
OmniSVG-8B + Ours
97.34
9.56
1.20
94.32
96.29
5.09
3.0k
98.38
8.47
1.12
94.25
98.47
4.99
2.3k
Appendix
Table A.2: Effect of Delta at 4B and 8B scale. Despite a significant margin between 4B and 8B base models, applying our method to both models produces similar results, motivating us to adopt a smaller variant.
Figure A.3: User study interface. Example of an interface with a question comparing Ours to Gemini 3 Flash.
Figure A.4: API evaluation prompts. System prompts used for evaluation with closed-source API models of (a) Image-to-SVG, (b) Image-to-TikZ.
DreamSim ↑
SigLIP ↑
CLIP ↑
LPIPS ↓
TED Norm ↓
C-BLEU ↑
Tokens ↓
w/ Outcome only
80.9
90.6
89.3
37.8
57.3
12.9
1.3k
w/ Process only
81.1
91.3
90.2
37.3
56.8
11.9
0.7k
w/ Process + Outcome (Ours)
83.8
92.5
91.2
35.4
55.5
14.4
1.0k
Appendix
Table B.1: Reward composition ablation on Image-to-TikZ. Ablation done at 30% of training schedule.
We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity image-to-SVG generation through self-correction. Existing MLLM-based approaches rely on single-pass open-loop inference, where the model receives visual input only once and must generate thousands of SVG code tokens without intermediate verification. This paradigm inevitably leads to geometric drift, error accumulation, and visual hallucination on complex images. RefineSVG overcomes this limitation by invoking an external rendering engine after an initial SVG generation pass to compare the rendered output against the target image. The comparison yields a multi-dimensional visual residual map (Diff-Map) that is fed back to the model as a ReAct-style correction signal, driving a targeted correction step. To support this render-observe-correct interaction, we further introduce an SVG-oriented semantic vocabulary that compresses token sequences by over 52%. A progressive training pipeline spanning supervised fine-tuning, rejection-sampling cold-start data construction, and end-to-end agentic reinforcement learning aligns the model with closed-loop visual correction. Extensive experiments show that RefineSVG consistently outperforms existing baselines in reconstruction fidelity, structural accuracy, and code efficiency.Code is available at https://github.com/liuxiaobo66/RefineSVG.
Shaobo Liu, Feiqiao Mao, Shuaishuai Zhou +4
Shenzhen University · Shenzhen, China · Peking University +1
Image-to-code generation tests whether a vision-language model (VLM) can recover the structure of an image enough to express it as executable code. Existing benchmarks either focus on narrow visual domains, depend on paired executable reference code, or rely on generic rubrics that miss domain-specific reconstruction errors. We introduce Vision2Code, a reference-code-free benchmark and evaluation framework for multi-domain image-to-code generation. Vision2Code contains 2,169 test examples from 15 source datasets that span charts and plots, geometry, graphs, scientific imagery, documents, and 3D spatial scenes. Models generate executable programs, which we render and score against the source image using a VLM rater with dataset-specific rubrics and deterministic guardrails for severe semantic failures. We report render-success diagnostics that separate code execution failures from reconstruction quality. Human validation shows that this evaluation protocol aligns better with human judgments than either a generic visual rubric or embedding-similarity baselines. Across nine open-weight and proprietary models, we find that image-to-code performance is domain-dependent: leading models perform well on regular chart- and graph-like visuals but remain weak on spatial scenes, chemistry, documents, and circuit-style diagrams. Finally, we show that evaluator-filtered model outputs can serve as training data to improve image-to-code capability, with Qwen3.5-9B improving from 1.60 to 1.86 on the benchmark without paired source programs. Vision2Code provides a reproducible testbed for measuring, diagnosing, and improving image-to-code generation. Our code and data are publicly available at https://image2code.github.io/vision2code/.
Reinforcement learning with verifiable rewards (RLVR) trains language models using programmatically checkable signals such as unit-test outcomes, enabling direct optimization for functional correctness in code generation. We conduct an empirical study of RLVR for Python code generation on the MBPP benchmark using two small models (Qwen3-0.6B and Llama3.2-1B) with LoRA fine-tuning. Across multiple reward formulations such as: unit-test-only rewards, static-analysis-only shaping via the Ruff linter, and a combined reward, we compare group-based policy optimization variants (GRPO and GSPO) and evaluate both functional correctness and behavioral diagnostics. In our experimental setting, RLVR improves pass@1 on MBPP test by up to 13 percentage points under proposed combined reward configuration. However, we find that reward shaping can induce systematic behavioral shifts: using only static-analysis penalties may bias the policy toward shorter completions that reduce lint errors without reliably improving functional correctness. In contrast, combined rewards mitigate this degeneration and yield more stable trade-offs between correctness and style constraints. Overall, our results highlight that RLVR effectiveness for code generation is highly sensitive to reward design and optimization granularity, and that diagnostics beyond pass@1, including generation length, Ruff severity profiles, and execution error types are useful for identifying failure modes.