cs.AISep 28, 2026

SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents

Authors: Bingqing Jiang, Guoxi Zhang, Jasper Wang, Auric Wang, Bingning Wang, Tianyi Lin, Zichao Yu, Yujin Han, +2 more

Organizations: The University of Hong Kong · WeChat, Tencent · Peking University · The Hong Kong Polytechnic University · City University of Hong Kong

Abstract

Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervision for intermediate decisions. Rubric-based rewards address this limitation through explicit intermediate criteria, but reliable rubrics are difficult to construct at scale and often disconnected from the procedure followed by the actor. We observe that a well-structured skill naturally specifies both how to act and what successful execution should achieve. Based on this insight, we introduce SkillRubric, which represents each skill through aligned actor-facing guidance and an evaluator-facing rubric. A multimodal verifier evaluates skill-defined goals using screenshots and tool outputs, assigning completion and progress rewards to the responsible turns. We further introduce an alternating co-evolution scheme that validates guidance revisions through paired rollouts under a frozen policy and rubric revisions offline under fixed guidance. Experiments across diverse multimodal agent benchmarks demonstrate consistent performance gains, while controlled paired rollouts further show that evolved skills provide more effective guidance for planning and tool use than their preceding versions.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

    Jul 2, 2026Jiayin Zhu, Kelong Mao, Yudong Guo +4SkillsRubrics

  2. SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution

    Sep 14, 2026Haoxiang Kang, Ming WenSkill EvolutionSkills

  3. SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing

    Jun 12, 2026Haowen Gao, Haoran Chen, Can Wang +5Skill EvolutionSkills