V-Gym: Enhancing Agentic Visual Reasoning via Skill-Data Co-Evolution
Authors: Bei Yan, Yuecong Min, Jie Zhang, Junqi Yang, Shiguang Shan, Xilin Chen
Organizations: State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China
Advances in multimodal understanding, reasoning, and tool use enable agents to tackle increasingly complex visual reasoning tasks. By distilling past execution experience into reusable skills, agents can transfer lessons from both successes and failures into future reasoning, reducing repeated errors and improving capabilities. However, limited experience may produce unreliable, poorly generalizable skills, while static datasets may lack the targeted and diverse practice needed for refinement. To address this gap, we introduce V-Gym, an autonomous framework that iteratively co-evolves procedural skills and multimodal practice data from execution trajectories. During skill evolution, V-Gym analyzes trajectories to distill and refine hierarchical skills, updating procedural guidance and applicability conditions while retaining an update only if it improves validation performance. During data evolution, V-Gym selects generation seeds by balancing data utility and exploration, then translates trajectory-identified bottlenecks into diverse, targeted practice data that expand the data bank after quality checks. The resulting practice outcomes feed back into subsequent skill updates, closing the loop for continual skill refinement. Experiments across diverse multimodal reasoning benchmarks show substantial improvements over baselines with multiple backbone models. Its evolved skills generalize across domains and models, while evolved data support more effective skill refinement, enabling autonomous diagnosis, targeted practice, and continual self-improvement.
Figures & tables
Figure 1: Shared trajectories for skill-data co-evolution. They guide skill refinement (top) and targeted data generation (bottom), updating both banks for later practice.
Figure 2: Overview of V-Gym. The agent uses the global skill and at most one routed task-specific skill in tool-assisted practice, recording trajectories and feedback. Skill evolution groups trajectories, consolidates patches by object, and accepts updates with positive independent validation gains. Data evolution ranks seeds by utility and exploration, then admits targeted generated data that pass self-checks. The updated banks support later practice.
Method
TIRBench
MMSearch-Plus
MMBrowseComp
Average
Avg.@3
Pass@3
Avg.@3
Pass@3
Avg.@3
Pass@3
Avg.@3
Pass@3
GPT-5.5
Baseline
40.9
57.7
20.3
36.0
15.6
22.7
25.6
38.8
Vanilla Tools
63.1
73.8
34.0
49.0
33.3
46.0
43.5
56.3
XSkill ( Jiang et al., 2026 )
64.7
77.4
34.3
54.0
36.2
54.7
45.1
62.0
Ace-Skill ( Xiong et al., 2026 )
64.7
78.7
36.0
50.0
36.7
56.0
45.8
61.6
Table 1: Task-solving performance across three backbones and three benchmarks (%).
Method
Target Benchmark
Avg.@3
Pass@3
TIRBench → VisualToolBench
Vanilla Tools
39.6
58.7
V-Gym (Ours)
46.2 ( ↑ 6.6)
66.7 ( ↑ 8.0)
MMSearch-Plus → AgentVista
Vanilla Tools
31.3
41.6
Table 2: Out-of-distribution transfer evaluation results (%). Skills acquired from the source domain or model are directly applied to the target domain or model for evaluation.
Figure 5
Figure 4: Skill evolution with different practice data on TIRBench with GPT-5.5. We compare (a) task-solving performance, (b) skill-update acceptance across epochs, and (c) example skills evolved with fixed data and our co-evolved data. Epoch 1 is shared warm-up.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Tool
Description
Main inputs
Shared by execution and generation
Web Search
Search the web via Serper API for titles, URLs, and text snippets.
Search related images via Serper image/Lens. ImgBB provides URLs for local reverse-image inputs.
search_type (str): text or reverse. query (str): text. image_url (str): reverse. max_results (int, optional).
Visit
Extract the main textual content of a webpage through the Jina Reader API.
url (str, required): page URL. goal (str): information to find.
Code Interpreter
Stateful Jupyter kernel for Python image processing (PIL/OpenCV), calculations, and data manipulation.
code (str, required): Python code.
Generation only
Appendix
Table 4: Tool capabilities and main inputs used in the experiments.
Dataset
Domain
Total
Train
Val.
Test
Sampling strategy
Visual Agentic Tool Use
TIRBench
Tool-Integrated Reasoning
1,215
195
630
390
Balanced random sampling across 13 task types.
VisualToolBench
Hybrid Tool Reasoning
1,204
–
–
150
Balanced random sampling of two single-turn types.
Multimodal Search
MMSearch-Plus
Multimodal Search
311
100
110
100
Random Sampling
MMBrowseComp
Multimodal Browsing
400
100
150
150
Random Sampling
Appendix
Table 5: Dataset domains, sizes, and partitions. All random sampling uses seed 42.
Dataset
Shared tools
Generation only
Code
Web
Image
Visit
Gen.
Edit
Capture
Visual Agentic Tool Use
TIRBench
✓
–
–
–
✓
✓
–
VisualToolBench
✓
✓
–
✓
–
–
–
Multimodal Search
MMSearch-Plus
✓
✓
✓
✓
–
–
✓
Appendix
Table 6: Tool availability by dataset. Shared tools are available for execution on each marked dataset and for generation on the three main datasets. Generation-only tools are used only in V-Gym data generation. Transfer targets have no target-domain generation.
Parameter
Value
Description
Models and execution settings
Solver model
Evaluated base model
Backbone used to solve tasks.
Judge model
GPT-5.5 (default)
Scores complete trajectories.
Embedding model
text-embedding-3-small
Similarity-based retrieval.
Test rollouts per instance
3
Independent test-time solutions.
Solver temperature
0.6
Practice, comparative validation, and test sampling.
Appendix
Table 7: Parameter settings for V-Gym.
Figure 5: Examples of unreliable answers or evidence, answer leakage or shortcuts, image-question mismatch, and Core-Challenge Drift, from left to right.
Figure 6: Examples of valid generated data. Generation aims to preserve core challenges and comparable difficulty while introducing meaningful variation.
Method
Mean SD (pp)
GPT-5.5
Gemini-3.5-Flash
Qwen-3.7-Flash
Baseline
3.16
3.27
1.66
Vanilla Tools
1.29
2.19
2.32
XSkill
2.46
—
—
Ace-Skill
2.23
—
—
SkillOPT
2.09
—
—
Appendix
Table 8: Mean per-benchmark sample SD of single-run accuracy (pp) for Table 1 . This differs from the SD of the across-benchmark mean. Dashes denote settings not evaluated.
Figure 7: Data-evolution analyses on TIRBench with GPT-5.5 at epoch 2. We compare (a) generation budgets with all selected seeds retained and (b) seed coverage at a fixed generation budget of ρ=0.6 .
School of Artificial Intelligence, Beihang University, Beijing, China · Hangzhou International Innovation Institute, Beihang University, Hangzhou, China · The Hong Kong Polytechnic University, Hong Kong SAR, China +3