Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
Figures & tables
Figure 1 : Co-evolution with adaptive visual environments and actors leads to better training.
Figure 2 : VICO turns self-evolution into an adaptive actor–environment loop. The EnvRewriter creates visual tasks, the Actor solves them, and shared rollouts update both policies: task rewards train the Actor, while difficulty-aligned rewards train the EnvRewriter.
Domain
Editing Interface
Label-Validity Rule
CLEVR [ 19 ]
Edit the scene graph: attributes, objects, shapes, or relations.
Re-execute the question program on the edited graph; reject rendering or ambiguity failures.
ChartQA [ 26 ]
Edit the CSV table under a two-tier rule: Tier 1 has a unique answer cell; Tier 2 has no unique cell mapping.
Table 1 : VICO keeps edited samples verifiable across domains. Structured domains use deterministic recomputation, while natural images use constrained answer preservation; neither needs judging or relabeling.
Baselines ( ↓ )
Mathematical Reasoning
General Vision-Centric
All
Datasets ( → )
MathVista
MathVision
ChartQA
MathVerse
WeMath
LogicVista
avg.
MMMU
MMVP
BLINK
avg.
avg.
Commercial VLMs (For reference)
GPT-4o
61.4
30.4
80.7
50.2
40.0
45.9
51.4
65.1
86.3
60.0
70.47
61.0
GPT-5
80.2
41.1
85.9
68.1
64.7
49.1
64.9
68.6
84.7
74.8
76.0
70.4
Gemini-2-Flash
73.4
41.3
77.0
54.4
57.1
56.2
56.9
52.6
83.0
59.6
65.1
62.5
Gemini-2.5-Pro
72.8
29.6
80.1
67.0
68.9
52.7
61.9
54.8
75.7
78.1
69.5
65.7
Table 2 : VICO achieves the strongest overall performance among comparable open-source VLMs. We report accuracy on nine multimodal reasoning benchmarks; bold marks the best result and underline marks the second-best among non-commercial models.
Baselines ( ↓ )
Seed Data
Mathematical Reasoning
Vision-Centric
Overall
Datasets ( → )
MathVista
MathVerse
ChartQA
WeMath
LogicVista
MMMU
BLINK
avg.
Chart-Specialized Methods
Chart-R1-7B [ 6 ]
∼ 33k
-
-
91.0
-
-
-
-
-
Bespoke-MiniChart-7B [ 28 ]
∼ 91k
-
-
89.5
-
-
-
-
-
ECD (Qwen2.5-VL-7B) [ 46 ]
∼ 320k
-
-
85.3
-
-
-
-
-
Qwen2.5-VL-7B Family
Table 3 : VICO improves VLM backbones with far less training data. The same pipeline improves Qwen2.5-VL-7B and InternVL3-8B, while using 16× – 160× less data than chart-specialized RL models. Overall average uses the seven displayed benchmarks; bold marks the best result and underline marks the second-best. *Average is computed over reported columns only.
Table 6
Figure 3 : VICO trains stably and induces an emergent curriculum. (a) GRPO reward improves smoothly. (b) ChartQA accuracy peaks at Round 4 with a +12.2 gain. (c) Actor responses grow by +49% , suggesting harder tasks and longer reasoning. (d) Gains concentrate on reasoning-heavy question types.
Baselines
ChartQA ( Δ )
Base (Qwen3-VL-8B)
76.9
Actor SFT only
79.4 +2.5↑
Co-evolution (w/ Vanilla Re-writer)
78.6 +1.7↑
Co-evolution (VICO)
89.1 +12.2↑
Table 6 : Rewriter initialization is not enough; co-evolution unlocks the full gain. Actor SFT (w/o co-training) and a vanilla EnvRewriter give modest gains ( +2.5% and +1.7% ), while an SFT-initialized EnvRewriter optimized in the actor–environment loop reaches +12.2% .
Figure 4 : The EnvRewriter calibrates bi-directional difficulty. On ChartQA, edits lower a strong Actor’s pass rate; on ZwZ, edits raise a weak Actor’s pass rate, moving both toward the sweet zone.
Baselines ( ↓ )
Mathematical Reasoning
Vision-Centric
Overall
MathVista
MathVision
ChartQA
MathVerse
WeMath
avg.
MMMU
MMVP
BLINK
avg.
avg.
Qwen3-VL-8B
67.7
31.5
76.9
50.6
72.0
59.8
55.8
77.7
68.9
67.5
63.6
Round1 (clean)
73.6
35.9
85.2
53.6
74.0
64.5
60.2
78.3
69.0
69.2
66.8
Round1 (noise-20%)
72.9
35.5
84.8
52.2
72.7
63.6
58.4
78.0
68.9
68.5
66.0
Table 7 : VICO remains robust under imperfect editing. Training with 20% failed-edit images causes only a slight drop, showing that the co-evolution loop does not need a flawless image editor.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Value
Actor optimizer
AdamW
Actor update
GRPO
Actor learning rate
1×10−6
Rollouts per sample
K=8 for pass-rate estimation
Sampling temperature
0.7
Top- p
1.0
Appendix
Table 8 : VICO uses one compact training recipe. Unless stated otherwise, the same core hyperparameters are used across visual domains.
Model & LoRA
Optimization
base model
Qwen3VL-8B-Instruct
learning rate
5.0×10−6
template
qwen3_vl
scheduler
cosine
finetuning type
LoRA
warmup ratio
0.1
LoRA rank
64
total epochs
10
LoRA alpha
128
per-device batch
2
LoRA target
all
grad. accum. steps
8
Appendix
Table 9 : RWR trains the EnvRewriter with pass-rate-weighted edit transitions. We train LoRA on Qwen3VL-8B-Instruct using LlamaFactory. Each retained edit transition is duplicated ⌊5⋅w(p^)⌋ times based on the Actor pass rate p^ on the sample produced by that edit turn.
Model & Reward
Rollout & Batch
base model
Qwen3VL-8B-Instruct
rollouts per prompt, N
8
algorithm
GRPO
train batch size
32
reward manager
binary
PPO mini-batch
8
KL in reward
False
PPO micro-batch/GPU
1
KL loss coeff.
0
log-prob micro-batch/GPU
1
entropy coeff.
0
rollout backend
vLLM
Appendix
Table 10 : GRPO trains the Actor with rollout-level verifiable rewards. We train Qwen3VL-8B-Instruct with verl and a vLLM rollout backend. We disable KL penalty in reward shaping and loss to avoid gradient interference, following [ 31 ] .
Arm
Editing interface
Answer rule
CLEVR
Scene graph edits over objects, attributes, and relations; rendered with Blender.
Recompute a′ by re-executing the question program.
SAM3 masks with protected target objects; edit only non-target regions.
Keep q′=q and a′=a under target-object protection.
Appendix
Table 11 : VICO uses three editing arms for verifiable visual variation. CLEVR recomputes answers from scene graphs, ChartQA uses a two-tier table rule, and ZwZ preserves answers by protecting target objects.
Figure 5 : CLEVR edits change difficulty while preserving verifiable labels. Each column shows one case before and after editing. The EnvRewriter modifies the scene graph, re-renders the image, and recomputes a′ by executing the original question program on the edited graph.
Figure 6 : ChartQA edits increase difficulty through table-level changes. Each column shows one case before and after editing. The EnvRewriter modifies the CSV-level representation and re-renders the chart. For Tier 1 examples, a′ is recovered by entity-anchored lookup in the edited table; for Tier 2 examples, CSV values are kept fixed and a′=a is preserved by construction.
Figure 7 : ZwZ edits preserve answers by protecting target objects. Green boxes mark answer-relevant targets, and blue boxes show other SAM3 regions. VICO keeps (q′=q,a′=a) fixed and edits only non-target regions, allowing background difficulty to change while preserving the original answer.
Figure 8 : Direct Nano Banana editing can break answer validity. Starting from the original image, direct generation can produce easier or harder variants, but the edits may change target relations or introduce ambiguous objects. As a result, the original question-answer pair may no longer be verifiable.
Figure 9 : VICO makes ZwZ easier while preserving the answer. The pipeline detects object regions, locks the answer-relevant targets, and edits only non-target regions. This changes visual difficulty while keeping the original spatial relation and answer verifiable.
Round
Sweet-zone hit rate (%)
R1
53.4
R2
29.5
R3
21.2
R4
20.5
R5
11.1
Appendix
Table 12: Sweet-zone hits decline as the Actor approaches saturation. The fraction of EnvRewriter trajectories in [0.35,0.65] falls monotonically as the Actor matures, matching the saturation regime predicted by Eq. E.6 .
Round
ChartQA (Qwen3-VL-8B)
Δ
Base
76.90
—
R1
85.24
8.34
R2
87.50
2.26
R3
88.47
0.97
R4
89.12
+0.65
R5
88.94
−0.18
Appendix
Table 13: ChartQA gains diminish as co-evolution saturates. Accuracy improves through R4, while round-over-round gain Δ shrinks and turns slightly negative at R5, consistent with the saturation regime predicted by our drift-corrected theory.
Figure 10 : EnvRewriter edit length stays stable. Multi-round training improves edit strength without making edit instructions longer.
Figure 11 : VICO improves accuracy across VLM backbones. Per-round ChartQA accuracy rises for both tested models, showing that the emergent curriculum is not tied to one model family.
Method
Data Cost
Training
Performance
Prepare Method
Num (RL)
Label Cost (Tokens)
Method
Interact
Time Cost
MathVerse
Mathvista
Qwen2.5-VL-7B
–
–
–
–
–
–
39.6
68.2
R1-OneVision-7B
Programmatic + human checks
10k
≥ 1.1 M
SFT+GRPO
✗
≥ 170 A100-Hours
46.4
64.1
VLAA-Thinker-7B
Teacher LLM-generated traces
25k
29.6 M
SFT+GRPO
✗
≥ 120 A100-Hours
48.2
68.0
OpenVLThinker-7B
Teacher LLM-generated solutions
9k
5.7 M
SFT+GRPO
✗
≥ 125 A100-Hours
47.9
70.2
MM-Eureka-Qwen-7B
Human-curated
15k
–
GRPO
✗
≈ 700 A100-Hours
50.3
73.0
Appendix
Table 14 : VICO cuts labeling cost while staying competitive. Label Cost counts teacher or judging LLM tokens with the Qwen2.5 tokenizer. Programmatic-label methods, including VICO, use no LLM labeling tokens; baseline training costs are estimated from reported RL samples, while VICO cost is directly measured.
Method
Labeled QA
ChartQA (%)
Chart-R1-7B [ 6 ]
∼ 33k
91.04
Bespoke-MiniChart-7B [ 28 ]
∼ 91k
89.5
ECD (Qwen2.5-VL-7B) [ 46 ]
∼ 320k
85.32
VICO-Qwen2.5-VL-7B (Ours)
2,000 seed inputs
87.96
VICO-Qwen3-VL-8B (Ours)
2,000 seed inputs
89.12
Appendix
Table 15 : VICO reaches competitive ChartQA accuracy with far fewer labeled QA pairs. All methods use ChartQA’s official relaxed accuracy with 5% numerical tolerance. VICO uses seed inputs and avoids extra labeled QA supervision, while chart-specialized baselines require tens or hundreds of thousands of labeled pairs.
Seed
ChartQA (%)
MathVista (%)
Seed 1 (reported)
85.2
73.8
Seed 2
84.5
72.9
Seed 3
85.2
71.7
Mean ± Std
85.0 ± 0.4
72.8 ± 1.1
Base (Qwen3VL-8B)
76.9
67.7
Gain over base
+8.1
+5.1
Appendix
Table 16 : VICO is stable across random seeds. We re-run the first CLEVR training round with three seeds and evaluate on ChartQA and MathVista. Standard deviation is far below the reported gains, so improvements are not driven by favorable seed selection.
Round
n
p^orig
p^final
Editing Intensity
ChartQA (strong Actor → EnvRewriter pushes down)
R1
2200
0.846
0.622
+0.225
R2
2200
0.829
0.502
+0.328
R3
2200
0.822
0.502
+0.321
R4
2200
0.848
0.387
+0.461
R5
2200
0.889
0.375
+0.514
Appendix
Table 17 : The EnvRewriter moves both strong and weak Actors toward the sweet zone. On ChartQA, it lowers the pass rate of a strong Actor; on ZwZ, it raises the pass rate of a weak Actor. Both converge toward [0.35,0.65] , showing actor-relative difficulty calibration.
Metric
Early round
Late round
Change
Editing intensity ∣p^orig−p^∣
0.22
0.51
+125%
EnvRewriter instruction length
973
981
+0.8%
Actor response length
811
1207
+49%
Appendix
Table 18 : Editing intensity grows without longer EnvRewriter instructions. Across rounds, edits become stronger while instruction length stays nearly fixed; Actor responses grow as tasks become harder.
Figure 12 : Saturation appears near the Actor’s capacity ceiling. Editing intensity grows by +125% , per-round gains shrink from +8.34% to −0.18% , and sweet-zone hit rate falls from 53.4% to 11.1% , matching Eq. E.6 .
Outcome
n
%
Both correct (Base ✓, R4 ✓)
1863
74.5
Improvement (Base ✗ → R4 ✓)
365
14.6
Regression (Base ✓ → R4 ✗)
64
2.6
Both wrong (Base ✗, R4 ✗)
208
8.3
Appendix
Table 19: Positive capability flips dominate regressions. On ChartQA, wrong-to-correct flips from Base to R4 outnumber correct-to-wrong flips by 5.7×1 , showing that VICO improves capability rather than merely moving errors across samples.
State
n
%
Already correct at Base
1927
77.1
Newly unlocked at R1
347
13.9
Newly unlocked at R2
37
1.5
Newly unlocked at R3
15
0.6
Newly unlocked at R4
3
0.1
Persistent failures
171
6.8
Appendix
Table 20 : Most new ChartQA capabilities unlock early and remain stable. Each sample is represented by a 5-bit correctness sequence across Base and R1–R4. The top panel shows first-correct rounds; the bottom lists common trajectories. The top five patterns cover 92.5% of samples, led by stable correctness ( 11111 ) and stable R1 unlocks ( 01111 ).
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4