cs.CVOct 8, 2026

ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

Authors: Hongxing Li, Dingming Li, Yixin Li, Yong Du, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, +1 more

Organizations: Zhejiang University

Abstract

Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored. We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents. Retrieved skills guide both inference and reward shaping, while successful trajectories are distilled back into the library, forming a closed feedback loop in which skill accumulation and policy improvement reinforce each other. An optional cold-start mechanism further accelerates early-stage learning. Evaluated on Sokoban, FrozenLake, and PrimitiveSkill, ViSkill achieves an overall success rate of 0.89, rising to 0.91 with cold-start initialization, outperforming all evaluated proprietary and open-source baselines while converging faster than standard PPO. Our code is available at https://github.com/ZJU-REAL/ViSkill.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation

    Jul 17, 2026Ruijie Sang, Yiqun Duan, Pinhan Fu +3Vision-Language NavigationObject-Goal Navigation

  2. EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

    Sep 1, 2026Wei Wang, Wenqiao Zhang, Yutong Lin +14Robotic ControlLarge Language Model-Based Robot Planning

  3. AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents

    May 18, 2026Pan Wang, Yihao Hu, Xiujin Liu +3Memory-Augmented VLMsReward Shaping