RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation
Organizations: National University of Singapore · Work done during an internship at Apple · Apple
Abstract
Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both challenges highlight limitations of relying solely on scalar rewards in multi-environment RL: they provide limited information about cross-environment relationships and no within-group reward contrast when rewards are identical. This motivates richer textual feedback, such as rubrics describing rollout behaviours, to guide learning. Beyond rubrics' usage as reward, we repurpose rubrics to guide both online data selection and policy supervision. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide the selection of data that aligns with the overall behavioural composition of the mixed-environment batch while limiting overlap with already-selected data. Available positive rubrics (describing desired behaviours) provide privileged context for an on-policy self-distillation teacher, supplying additional token-level supervision, while negative rubrics (describing undesired behaviours) guide subsequent rollout generation away from recurring failure modes. Together, these components form RISED. Across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. Rubric-based analysis of RISED can further characterize the behavioural changes accompanying these gains.
Figures & tables
| Rubric-level | Skill-level | Gradient | ||||||
|---|---|---|---|---|---|---|---|---|
| Val. envs | pos | neg | all | pos | neg | all | full layer. | |
| Held-out | ||||||||
| All env | ||||||||
| Model | Method | ALFWorld | WebShop | DBBench | Overall |
|---|---|---|---|---|---|
| Qwen2.5-3B | GRPO-64 | 0.681 0.041 | 0.229 0.004 | 0.541 0.013 | 0.484 0.015 |
| GRPO-128 | 0.748 0.000 | 0.244 0.009 | 0.552 0.009 | 0.515 0.006 | |
| Global Variance | 0.771 0.020 | 0.248 0.013 | 0.550 0.008 | 0.523 0.006 | |
| Round-robin Variance | 0.805 0.005 | 0.226 0.005 | 0.557 0.011 | 0.530 0.003 | |
| All-Rubric SD | 0.747 0.003 | 0.238 0.013 | 0.557 0.007 | 0.514 0.008 | |
| RISED | 0.800 0.023 | 0.256 0.003 | 0.569 0.002 | 0.542 0.008 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Environment | Rubric | Illustrative behaviour |
|---|---|---|
| WebShop | formulates_constrained_search_query | Constructs a search query using relevant task constraints. |
| ALFWorld | systematic_object_search | Searches plausible locations systematically to find a target object. |
| DBBench | explores_database_schema | Inspects the database structure to identify relevant tables and columns. |
| Skill | Rubric | Illustrative behaviour |
|---|---|---|
| ALFWorld: taking an apple fails because the fridge is closed. | ||
| error_recovery_and_exploration | revises_plan_after_failure | Opens the fridge before retrying the action. |
| error_recovery_and_exploration | repeats_failed_action | Repeats the failed action without opening the fridge. |
| ALFWorld: the task requires placing a hot mug on a side table. | ||
| object_manipulation | uses_appliance_for_state_change | Heats the mug using a microwave before placing it on the table. |
| instruction_following_and_grounding | ignores_task_constraint | Places the mug on the table without first heating it. |
| Environment | Train | Validation | Training Data / Selection | Validation Data / Selection |
|---|---|---|---|---|
| ALFWorld | 3,150 | 109 | AG train_valid | AG new_std |
| DBBench | 4,803 | 300 | AG db_out_new.jsonl | AG standard.jsonl |
| WebShop | 11,000 | 200 | AG indices | AG indices |
| AppWorld | 72 | 57 | Easy and medium training tasks | Complete dev split |
| Qwen2.5-3B-Instruct | Qwen3-4B | |
| Rollout | ||
| Sampling temperature | ||
| Max new tokens per turn | ||
| Reasoning mode | — | disabled |
| Max turns per episode | ||
| Max episode length (tokens) | ||
| Rubric-level | Skill-level | Gradient | ||||||
|---|---|---|---|---|---|---|---|---|
| Val. envs | pos | neg | all | pos | neg | all | full layer. | |
| Held-out | ||||||||
| All env | ||||||||
| Signal agreement | Held-out | All env | |||||
|---|---|---|---|---|---|---|---|
| Feature | mean | run 1 | run 2 | run 1 | run 2 | max | |
| Rubric, all | |||||||
| Rubric, violates | |||||||
| Rubric, demonstrates | |||||||
| Skill, all | |||||||
| Skill, violates | |||||||
| Method | Roll. wait | Tag-wait | Ref | Teacher | Actor | Other | Step | Relative |
|---|---|---|---|---|---|---|---|---|
| GRPO-64 | 0.5 | – | 25.7 | – | 89.7 | 39.6 | 155.5 | |
| GRPO-128 | 0.8 | – | 48.0 | – | 167.0 | 33.8 | 249.6 | |
| Global Variance | 16.5 | – | 27.9 | – | 95.6 | 39.1 | 179.1 | |
| Round-robin Variance | 21.9 | – | 26.5 | – | 88.4 | 40.5 | 177.3 | |
| RISED | 8.6 | 63.0 | 27.2 | 30.6 | 92.1 | 31.4 | 252.9 |