cs.CVOct 7, 2026

VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

Authors: Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, +3 more

Organizations: Virginia Tech · Cisco · NVIDIA · UChicago · Georgia Tech · Abaka AI · UC Berkeley · UGA · WashU

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

    Jul 22, 2026Md Tanvirul AlamSynthetic Data GenerationVLM Reasoning

  2. MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

    May 2, 2026Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5Vision-Language ModelsReinforcement Learning

  3. DUEL: Adversarial Self-Play for Multimodal Reasoning

    May 24, 2026Lin Qiu, Hanqing Zeng, Yao Liu +3VLM RobustnessVLM Reasoning