SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents
Organizations: The University of Hong Kong · WeChat, Tencent · Peking University · The Hong Kong Polytechnic University · City University of Hong Kong
Abstract
Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervision for intermediate decisions. Rubric-based rewards address this limitation through explicit intermediate criteria, but reliable rubrics are difficult to construct at scale and often disconnected from the procedure followed by the actor. We observe that a well-structured skill naturally specifies both how to act and what successful execution should achieve. Based on this insight, we introduce SkillRubric, which represents each skill through aligned actor-facing guidance and an evaluator-facing rubric. A multimodal verifier evaluates skill-defined goals using screenshots and tool outputs, assigning completion and progress rewards to the responsible turns. We further introduce an alternating co-evolution scheme that validates guidance revisions through paired rollouts under a frozen policy and rubric revisions offline under fixed guidance. Experiments across diverse multimodal agent benchmarks demonstrate consistent performance gains, while controlled paired rollouts further show that evolved skills provide more effective guidance for planning and tool use than their preceding versions.
Figures & tables
| Method | IID Benchmarks | OOD Benchmarks | Avg. | ||||||
| AgentVista | BrowseComp-VL | VDR-Bench | VisualToolBench | VisBrowse | MMSearch+ | MMSearchExam | TIR-Bench | ||
| Closed-source model references | |||||||||
| Gemini 3.1 Flash-Lite | 26.61 | 43.81 | 20.00 | 31.01 | 33.14 | 33.76 | 26.15 | 27.00 | 30.19 |
| GPT-5 nano | 12.84 | 35.79 | 5.67 | 14.95 | 24.85 | 5.14 | 24.73 | 17.50 | 17.68 |
| GPT-5 mini | 16.51 | 45.15 | 11.67 | 27.54 | 36.09 | 14.79 | 34.98 | 24.50 | 26.40 |
| Open-source model references | |||||||||
| Backbone | Variant | Components | Average Success Rate (%) | ||||
| Outcome RL | Rubric Reward | Co-Evolution | IID | OOD | Overall | ||
| Qwen3-VL-8B | SFT-only | – | – | – | 11.81 | 7.59 | 9.70 |
| GRPO-only | – | – | 13.38 | 9.70 | 11.54 | ||
| GRPO-Rubric | – | 15.46 | 11.70 | 13.58 | |||
| Ours | 19.91 | 15.86 | 17.88 | ||||
| Qwen3.5-9B | SFT-only | – | – | – | 21.77 | 21.60 | 21.68 |
| Backbone | Variant | IID | OOD | Overall | |
| Qwen3-VL-8B | Ours | 5 | 16.34 | 12.79 | 14.57 |
| Ours | 10 | 19.91 | 15.86 | 17.88 | |
| Ours | 20 | 18.35 | 13.98 | 16.17 | |
| GRPO-Rubric | 15.46 | 11.70 | 13.58 | ||
| Guidance evolution only | 10 | 18.68 | 15.04 | 16.86 | |
| Rubric evolution only | 10 | 17.81 | 14.84 | 16.33 |
| Model | IID Benchmarks | OOD Benchmarks | Avg. | ||||||
| AgentVista | BrowseComp-VL | VDR-Bench | VisualToolBench | VisBrowse | MMSearch+ | MMSearchExam | TIR-Bench | ||
| Qwen3-VL-8B-Instruct | 6.7 | 4.6 | 3.9 | 8.4 | 5.8 | 5.0 | 4.6 | 7.9 | 5.9 |
| Ours (Qwen3-VL-8B, w/ skill) | 5.9 | 3.3 | 3.1 | 6.1 | 5.2 | 3.2 | 3.1 | 6.0 | 4.5 |
| Qwen3.5-9B | 10.7 | 6.8 | 9.6 | 9.5 | 11.3 | 9.7 | 9.6 | 8.9 | 9.5 |
| Ours (Qwen3.5-9B, w/ skill) | 4.8 | 4.7 | 5.7 | 3.1 | 6.4 | 4.6 | 5.5 | 1.6 | 4.6 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Domain | Setting | Available | Train | Eval |
| AgentVista | Comprehensive Multimodal Tool Use | IID | 209 | 100 | 109 |
| BrowseComp-VL | Multimodal Browsing | IID | 399 | 100 | 299 |
| VDR-Bench | Multimodal Deep Research | IID | 2,000 | 100 | 300 |
| VisualToolBench | Tool-Assisted Visual Reasoning | IID | 1,204 | 100 | 200 |
| VisBrowse-Bench | Visual-Native Browsing | OOD | 169 | – | 169 |
| MMSearch-Plus | Multimodal Search | OOD | 311 | – | 311 |
| Tool | Function | Backend | Call limit |
| Zoom | Crops or enlarges selected image regions for fine-grained visual inspection. | Local image processor | No separate limit |
| Code interpreter | Executes Python code for image processing, numerical computation, and data analysis. | Isolated code sandbox | No separate limit |
| Web search | Retrieves relevant webpages and textual search results for a text query. | SerpAPI | 7 per trajectory |
| Image search | Retrieves visually related images and their associated webpages and sources. | SerpAPI; ImgBB for temporary hosting | 5 per trajectory |
| Visit | Opens a specified webpage and extracts its textual and visual content. | Jina and Firecrawl | No separate limit |
| Group | Setting | Value |
| Inference | Serving backend | SGLang (OpenAI-compatible) |
| Sampling temperature | 0.2 | |
| Top- | 1.0 | |
| Rollouts per task | 1 | |
| Maximum interaction turns | 20 | |
| Maximum context length | 131,072 tokens |