Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both challenges highlight limitations of relying solely on scalar rewards in multi-environment RL: they provide limited information about cross-environment relationships and no within-group reward contrast when rewards are identical. This motivates richer textual feedback, such as rubrics describing rollout behaviours, to guide learning. Beyond rubrics' usage as reward, we repurpose rubrics to guide both online data selection and policy supervision. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide the selection of data that aligns with the overall behavioural composition of the mixed-environment batch while limiting overlap with already-selected data. Available positive rubrics (describing desired behaviours) provide privileged context for an on-policy self-distillation teacher, supplying additional token-level supervision, while negative rubrics (describing undesired behaviours) guide subsequent rollout generation away from recurring failure modes. Together, these components form RISED. Across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. Rubric-based analysis of RISED can further characterize the behavioural changes accompanying these gains.
Figures & tables
Figure 1 : Overview of RISED . Sec. 2.2 describes the offline rubric discovery. Sec. 2.4 and Sec. 3 describe the selection and distillation, respectively. Together, these form RISED , which uses rubrics to select training data across diverse agentic environments and reuses them for on-policy self-distillation.
Rubric-level
Skill-level
Gradient
Val. envs
pos
neg
all
pos
neg
all
full layer.
Held-out
ρ
−0.52
+0.04
+0.60
−0.69
+0.30
+0.43
+0.18
r
−0.49
+0.07
+0.77
−0.61
+0.43
+0.52
+0.21
All env
ρ
+0.32
+0.57
+0.81
+0.25
+0.69
+0.75
+0.28
r
+0.76
+0.84
+0.87
+0.69
+0.87
+0.87
+0.27
Table 1: Correlation between similarity signals and reward improvement under WebShop-only training. Spearman ( ρ ) and Pearson ( r ) correlations are pooled over checkpoint–environment pairs across steps. Held-out results exclude WebShop. Positive-only, negative-only, and combined profiles are denoted by pos, neg, and all, respectively.
Figure 2 : Validation dynamics for RISED compared with baseline . Validation pass rate during joint training on ALFWorld, WebShop, and DBBench with Qwen2.5-3B-Instruct , shown per environment and overall. RISED achieves higher validation pass rates than the baselines, both overall and in each environment, during most of the middle and later stages of training.
Model
Method
ALFWorld
WebShop
DBBench
Overall
Qwen2.5-3B
GRPO-64
0.681 ± 0.041
0.229 ± 0.004
0.541 ± 0.013
0.484 ± 0.015
GRPO-128
0.748 ± 0.000
0.244 ± 0.009
0.552 ± 0.009
0.515 ± 0.006
Global Variance
0.771 ± 0.020
0.248 ± 0.013
0.550 ± 0.008
0.523 ± 0.006
Round-robin Variance
0.805 ± 0.005
0.226 ± 0.005
0.557 ± 0.011
0.530 ± 0.003
All-Rubric SD
0.747 ± 0.003
0.238 ± 0.013
0.557 ± 0.007
0.514 ± 0.008
RISED
0.800 ± 0.023
0.256 ± 0.003
0.569 ± 0.002
0.542 ± 0.008
Table 2 : Main results of RISED compared with baselines. Final-checkpoint pass rates on ALFWorld, WebShop, and DBBench, and overall, for RISED and five baselines on two backbones. Bold and underlined values indicate the best and second-best results, respectively, within each backbone. RISED achieves the highest mean overall pass rate and ranks among the top two methods in every individual environment.
Figure 3 : Component ablations. Validation pass rate across training steps for RISED (pink), RISED without selection (green), RISED without distillation (purple). Combining selection and distillation yields the strongest overall performance, while their individual contributions vary across environments.
Figure 4 : Positive-rubric accumulation and task performance. Solid lines show cumulative counts of training rollouts exhibiting the indicated rubric (left axis); dashed lines show validation pass rates (right axis). Pink denotes RISED and blue denotes GRPO-64. RISED exhibits positive behaviours more frequently than GRPO-64.
Figure 5 : Joint training with a challenging fourth environment. Validation pass rates for RISED (pink) and GRPO-64 (blue) when AppWorld is added. RISED achieves higher overall performance and shows emerging progress on AppWorld, while GRPO-64 remains near zero for AppWorld.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Environment
Rubric
Illustrative behaviour
WebShop
formulates_constrained_search_query
Constructs a search query using relevant task constraints.
ALFWorld
systematic_object_search
Searches plausible locations systematically to find a target object.
DBBench
explores_database_schema
Inspects the database structure to identify relevant tables and columns.
Appendix
Table 3: Example rubrics grouped under the shared skill search_and_navigation . Behaviour summaries are illustrative.
Skill
Rubric
Illustrative behaviour
ALFWorld: taking an apple fails because the fridge is closed.
error_recovery_and_exploration
revises_plan_after_failure
Opens the fridge before retrying the action.
error_recovery_and_exploration
repeats_failed_action
Repeats the failed action without opening the fridge.
ALFWorld: the task requires placing a hot mug on a side table.
object_manipulation
uses_appliance_for_state_change
Heats the mug using a microwave before placing it on the table.
instruction_following_and_grounding
ignores_task_constraint
Places the mug on the table without first heating it.
Appendix
Table 4: Positive and negative rubric examples for two ALFWorld tasks. Within each task, the first row illustrates a positive rubric and the second a negative rubric. Actions are illustrative rather than verbatim rollout excerpts.
Environment
Train
Validation
Training Data / Selection
Validation Data / Selection
ALFWorld
3,150
109
AG train_valid
AG new_std
DBBench
4,803
300
AG db_out_new.jsonl
AG standard.jsonl
WebShop
11,000
200
AG indices [1000,12000)
AG indices [0,200)
AppWorld
72
57
Easy and medium training tasks
Complete dev split
Appendix
Table 5: Training and validation data for the four environments. Sizes count task instances, not generated rollouts. AppWorld training includes only easy and medium tasks, whereas its validation set includes all difficulty levels. The main experiments and ablations use the first three environments; the preliminary correlation analysis and additional experiments use all four. AG denotes AgentBench
Qwen2.5-3B-Instruct
Qwen3-4B
Rollout
Sampling temperature
0.8
0.6
Max new tokens per turn
1024
2048
Reasoning mode
—
disabled
Max turns per episode
20
Max episode length (tokens)
8192
Appendix
Table 6: Training configuration for the two backbones. Values shared by both are given once; where they differ, both are listed.
Rubric-level
Skill-level
Gradient
Val. envs
pos
neg
all
pos
neg
all
full layer.
Held-out
ρ
+0.50
+0.32
+0.47
+0.46
+0.40
+0.40
+0.20
r
+0.54
+0.42
+0.43
+0.51
+0.36
+0.38
+0.18
All env
ρ
+0.77
+0.69
+0.76
+0.74
+0.73
+0.72
+0.41
r
+0.80
+0.89
+0.92
+0.63
+0.68
+0.79
+0.41
Appendix
Table 7 : Correlation between similarity signals and reward improvement under DBBench-only training. Spearman ( ρ ) and Pearson ( r ) correlations are pooled over checkpoint–environment pairs at steps 50–500, relative to the step-25 baseline. Held-out results exclude DBBench. Positive-only, negative-only, and combined profiles are denoted by pos, neg, and all, respectively.
Signal agreement
Held-out ρ
All env ρ
Feature
r
mean ∣Δ∣
run 1
run 2
run 1
run 2
max ∣Δρ∣
Rubric, all
0.9992
0.0115
+0.595
+0.588
+0.814
+0.809
0.007
Rubric, violates
0.9959
0.0260
+0.040
+0.052
+0.570
+0.576
0.012
Rubric, demonstrates
0.9663
0.0502
−0.522
−0.474
+0.320
+0.345
0.048
Skill, all
0.9986
0.0142
+0.431
+0.426
+0.750
+0.746
0.005
Skill, violates
0.9919
0.0312
+0.302
+0.362
+0.688
+0.712
0.060
Appendix
Table 8: Tagger reproducibility across two runs. Signal agreement is measured over checkpoint–environment pairs using Pearson correlation ( r ) and mean absolute difference (MAD). Transfer correlations are pooled Spearman coefficients ( ρ ), shown as Run 1 / Run 2.
Figure 6 : Rubric roles in teacher conditioning on Qwen2.5-3B-Instruct. We show positive-only conditioning without negative prompt injection (brown) and all-rubric conditioning with selection (green). Full RISED (pink) is included for reference. Positive-only conditioning outperforms all-rubric conditioning overall, particularly on ALFWorld, while full RISED learns faster initially than the positive-only variant, most clearly on DBBench.
Figure 7 : Negative rubrics reveal distinct stages of behavioural improvement in DBBench. Curves track the fraction of rollouts exhibiting each undesired behaviour. Markers indicate when each undesired behaviour rate suddenly falls to half its initial level. Repeated failed actions and false completion claims reach this milestone substantially earlier than hallucinated observations or actions.
Figure 8 : Effect of rubric judge size. Validation pass rate for RISED with 8B, 14B, and 32B judges. The 14B judge remains competitive, while the 32B judge achieves the best overall performance.
Method
Roll. wait
Tag-wait
Ref
Teacher
Actor
Other
Step
Relative
GRPO-64
0.5
–
25.7
–
89.7
39.6
155.5
1.00×
GRPO-128
0.8
–
48.0
–
167.0
33.8
249.6
1.61×
Global Variance
16.5
–
27.9
–
95.6
39.1
179.1
1.15×
Round-robin Variance
21.9
–
26.5
–
88.4
40.5
177.3
1.14×
RISED
8.6
63.0
27.2
30.6
92.1
31.4
252.9
1.63×
Appendix
Table 9: Per-step wall-clock time (seconds) for Qwen2.5-3B-Instruct across three environments using 32 H100 GPUs for training and rollout generation. Methods requiring rubric annotations additionally use a dedicated eight-GPU judge server. Reported values are averaged over first 300 steps. Roll. wait measures blocking on batch retrieval; Tag-wait measures blocking to complete rubric annotations. Ref computes reference-policy log probabilities for sampled rollouts; Teacher computes rubric-conditioned teacher log probabilities for distillation. Actor includes the policy forward pass, backpropagation, and optimizer update. Other is the remaining step time, including checkpointing and trainer overhead.
Reinforcement learning pipelines for Large Language Model (LLM) training often rely on manually redesigned environments between stages, requiring practitioners to heuristically infer which configuration will best improve the current policy. To automate this process, we propose the LLM-as-Environment-Engineer framework in which the current policy model analyzes failure trajectories together with contextual information and proposes modifications to the next-stage training environment configuration. We also introduce MAPF-FrozenLake, a controllable testbed whose generator exposes multi-dimensional environment configurations, making it suitable for studying and benchmarking environment redesign. On this testbed, we condition the environment engineer on structured summaries of policy behavior, failure cases, and environment statistics, from which it produces the configuration for the next training stage. With Qwen3-4B as the backbone, our framework achieves the strongest aggregate performance on our benchmarks, outperforming larger proprietary LLMs (e.g., GPT, Gemini) and fixed-environment training baselines. We further analyze which forms of context are most effective, finding that successful environment updates rely on failure evidence and preserve configurations that already work. Interestingly, the current RL checkpoint serves as a better environment engineer than the original base model, suggesting that policy learning improves the model's ability to diagnose its remaining weaknesses.
Chao Chen, Chengzu Li, Zhiwei Li +2
1LARK, HKUST (GZ) · University of Cambridge · 3HKUST
Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.
Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation trajectory. Existing rubric-based reinforcement learning (RL) methods compress fine-grained criterion-level feedback into scalar rewards, making persistent capability gaps difficult to target under limited on-policy exploration. We propose RISE-RL (Rubric-Informed Selective Exploration), which uses repeatedly missed rubric criteria to elicit privileged trajectories that are difficult to discover through unguided exploration alone. RISE-RL retains only trajectories whose complete-rubric reward exceeds the mean reward of natural rollouts, and then re-evaluates them under the original prompt to emphasize behaviors that remain weakly supported by the natural policy. The resulting guidance signal is optimized through a separate auxiliary objective and removed once its additional benefit diminishes. Experiments with 4B and 14B models across writing, chat, health, and science show that RISE-RL achieves the highest mean score on every evaluated benchmark under guidance-free evaluation. Compared with standard Rubric-RL, it improves the average score by 1.3 points at the 4B scale and 3.3 points at the 14B scale, including a 6.0-point gain on CreativeWriting-V3. It also improves creative-writing diversity and yields gains on objectively scored medical and scientific benchmarks. These results indicate that selective internalization through reward filtering and policy support shaping is effective for open-ended reinforcement learning.
Jinkun Hou, Zhuo Liu, Huimin Ren +3
1Peking University · 2Beijing Institute of Technology · 3Li Auto Inc.