Reinforcement learning is increasingly used to post-train vision-language models for image-to-code generation, such as generating SVG code from a reference image, by optimizing rewards computed from the final rendered output. However, relying on a single terminal reward provides sparse feedback that is poorly aligned with the contribution of individual tokens. A generated program may contain operations that accurately reproduce some parts of the target image alongside others that introduce errors, yet all tokens are trained from the same final outcome. We observe that many intermediate code prefixes are not only executable, but already produce meaningful partial renders that reflect progress toward the target. This property provides a natural source of denser supervision during generation. Based on this observation, we introduce IR4RL, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence. We evaluate our approach on Image-to-SVG and Image-to-TikZ generation. Across both tasks, our method improves over supervised fine-tuning and standard GRPO, yielding new state-of-the-art open-source models. This shows that intermediate rendering provides a simple and effective source of process supervision for RL post-training of image-to-code models.
Figures & tables
Figure 1: We introduce a method that leverages I ntermediate R enders for R einforcement L earning ( IR4RL ) post-training of image-to-code VLMs as a supervision signal, allowing us to improve over outcome-only RL and achieve state-of-the-art results in Image-to-SVG and Image-to-TikZ tasks.
Figure 2: Process Reward from Intermediate Renders. Rollouts 1 and 2 reach similar final quality through different generation trajectories while Rollout 3 makes early progress that later operations reverse. Outcome-only advantage rewards all tokens within a rollout equally without distinguishing helpful from harmful decisions. Our method complements it with localized feedback leveraging changes in visual score between consecutive renders (the arrows Δj ).
Figure 3: Qualitative evaluation of Image-to-SVG. We compare LIVE, Gemini-3-Flash, task-specific SVG models, and post-training baselines. Our method better preserves structure, color, geometry, and fine details, while competing methods more often omit or distort visual elements.
MMSVGBench-Illustrations
MMSVGBench-Icons
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
Optimization-based
DiffVG
92.03
9.23
0.35
94.10
93.52
4.85
79.6k
90.97
9.24
0.44
93.76
95.29
4.89
79.5k
LIVE
94.55
10.02
0.72
95.48
93.68
4.99
8.4k
94.24
9.18
0.86
95.19
95.60
4.91
8.4k
General-purpose (M)LLMs
Qwen3-VL-235B
92.81
28.32
5.30
87.89
89.82
4.82
5.0k
92.23
29.90
7.78
84.77
91.61
4.77
5.9k
Table 1: Quantitative evaluation on MMSVGBench. We compare optimization-based methods, general-purpose VLMs, SVG models, and post-training baselines on the Illustrations and Icons splits. Our method surpasses baselines, substantially improves the SFT base model, and performs best when process and outcome supervision are combined. Bold and underline denote the best and second-best results among learned methods, excluding optimization-based methods.
Gemini 3 Flash
InternSVG-8B
OmniSVG-4B (SFT)
OmniSVG-4B + RAFT
OmniSVG-4B + Outcome RL
Users ↑
60.6%
76.3%
92.7%
74.9%
74.0%
VLM ↑
54.0%
74.0%
94.0%
78.0%
76.0%
Table 2: User Study and VLM evaluation. 2AFC win rate (%) of our method vs. each baseline demonstrates our method is preferred by humans and a VLM (VLM-human agreement: 88%).
Figure 5: Qualitative reward composition ablation. Process-only training recovers much of the gain over outcome-only, while combining both rewards yields the most faithful reconstructions.
DreamSim ↑
SigLIP ↑
CLIP ↑
LPIPS ↓
KID ↓
C-BLEU ↑
TED ↓
Tokens ↓
General-purpose (M)LLMs
Qwen3-VL-235B
80.3
91.5
89.0
40.4
0.74
3.4
54.1
3.9k
Gemini 3 Flash
91.0
96.1
94.3
27.2
-0.05
6.6
51.7
2.4k
Sonnet 5
85.4
94.0
92.2
35.9
0.06
3.7
52.8
0.7k
GPT-5.2
84.9
93.8
91.4
35.7
0.41
4.4
54.0
1.5k
TikZ VLMs
Table 3: Quantitative evaluation on DaTikZ-v3. Our method improves DeTikZify-v2 and post-training baselines across the visual reconstruction metrics while producing substantially shorter programs. Bold and underline mark best and second-best results among task-specific TikZ models.
Figure 6: Ablations of process supervision. (a) Increasing the process contribution improves performance up to α=10 . (b) Reward propagation performs best at λ=0.9 . (c) More frequent intermediate renders consistently improve performance. Stars mark our default settings.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1: Best-of- K sampling on SVG generation. For each image, we sample K candidate SVGs and report the mean reward of the best candidate. Across multiple datasets and temperature values, our method performs best across all sampling budgets, with the largest gains at small K .
Figure A.2: Prefix closure for (a) Image-to-SVG, (b) Image-to-TikZ. A model generates a program code y that is separated into granular segments y1:bj , well-formed by a closure operator C into Pj=C(y1:bj) and rendered. This process depends on the task and a tokenizer. Specifically, OmniSVG does not generate the <svg> tag, while in DeTikZify the opening code is part of the sequence.
MMSVGBench-Illustrations
MMSVGBench-Icons
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
700 samples
97.48
9.79
1.17
94.22
96.05
5.09
2.5k
98.26
8.67
1.23
93.88
98.01
4.98
2.1k
10k samples
97.26
10.11
1.27
94.11
95.35
5.08
2.2k
98.49
8.42
1.20
94.12
98.35
5.02
1.9k
Appendix
Table A.1: Effect of the number of training samples. We observe that increasing the size of the trainset provides only marginal gains on Icons subset at the cost of mild degradation on Illustrations, and increases the overall generation length. We thus opt for a smaller trainset in all experiments.
MMSVGBench-Illustrations
MMSVGBench-Icons
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
DINO ↑
LPIPS ↓
MSE ↓
SSIM ↑
CLIP ↑
Aesthetic ↑
Tokens ↓
OmniSVG-4B
85.48
22.26
5.11
89.70
82.03
4.55
11.3k
89.21
19.82
5.80
87.96
88.42
4.65
8.4k
OmniSVG-8B
88.81
21.15
4.98
88.85
85.57
4.68
9.1k
92.04
18.09
4.92
89.28
91.89
4.79
6.0k
OmniSVG-4B + Ours
97.48
9.79
1.17
94.22
96.05
5.09
2.5k
98.26
8.67
1.23
93.88
98.01
4.98
2.1k
OmniSVG-8B + Ours
97.34
9.56
1.20
94.32
96.29
5.09
3.0k
98.38
8.47
1.12
94.25
98.47
4.99
2.3k
Appendix
Table A.2: Effect of Delta at 4B and 8B scale. Despite a significant margin between 4B and 8B base models, applying our method to both models produces similar results, motivating us to adopt a smaller variant.
Figure A.3: User study interface. Example of an interface with a question comparing Ours to Gemini 3 Flash.
Figure A.4: API evaluation prompts. System prompts used for evaluation with closed-source API models of (a) Image-to-SVG, (b) Image-to-TikZ.
DreamSim ↑
SigLIP ↑
CLIP ↑
LPIPS ↓
TED Norm ↓
C-BLEU ↑
Tokens ↓
w/ Outcome only
80.9
90.6
89.3
37.8
57.3
12.9
1.3k
w/ Process only
81.1
91.3
90.2
37.3
56.8
11.9
0.7k
w/ Process + Outcome (Ours)
83.8
92.5
91.2
35.4
55.5
14.4
1.0k
Appendix
Table B.1: Reward composition ablation on Image-to-TikZ. Ablation done at 30% of training schedule.