As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric-Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency. The project page is at https://foundation-model-research.github.io/ARISE .
Figures & tables
Figure 1: Performance and motivation of Arise . (a) SkillsBench v1.1 pass rates versus model size, highlighting the improvement over Qwen3.5-27B and competitive performance with substantially larger models. (b) Fixed criteria can miss emerging capability gaps, while sampled rollouts may all fail despite making partial progress. Arise uses rollout evidence to adapt rubrics, skill guidance, and task selection as agent capabilities evolve.
Figure 2: Overview of Arise . Tasks sampled via capability-based priorities generate rollouts with selectively activated skill guidance. Rubric evaluations provide behavioral feedback for policy optimization and, alongside rollout evidence, drive rubric–skill evolution and subsequent task sampling.
SkillsBench v1.1
TB v2.1
Model
SE
NS
OW
IP
FE
MR
CS
MC
Overall
Overall
Proprietary Models
GPT-5.4 Mini
27.1
35.7
45.2
21.4
25.9
41.7
33.3
66.7
34.5
59.2
GPT-5.5
63.4
77.9
76.2
57.5
37.0
95.0
69.0
60.0
67.3
84.3
Claude Opus-4.7
58.3
83.3
54.8
54.8
44.4
50.0
57.1
53.3
58.6
83.1
Open-Weight Models
Table 1: Pass rates (%) on SkillsBench v1.1 and Terminal-Bench v2.1 (TB). SkillsBench categories are software engineering (SE), natural science (NS), office and white collar (OW), industrial and physical systems (IP), finance and economics (FE), mathematics and operations research/formal reasoning (MR), cybersecurity (CS), and media and content production (MC). For SkillsBench, GPT-5.5 results are from the official leaderboard; dashes denote models without reported results from either that leaderboard or our own evaluation. Bold values indicate the best results within comparable-scale models.
Figure 3: Ablation results and training efficiency on SkillsBench v1.1. (a) Pass rates for Arise and its ablated variants. (b) Pass rates over training under the same per-step rollout budget.
Figure 4: Behavioral feedback and exploration during training. (a) Active, added, and retired rubric counts. (b) Mean rubric reward and verifier-based task pass rate in 10-step windows. (c) All-failure rollout group rates, averaged equally across tasks in the same fixed task set.
Figure 5: Adaptive sampling probabilities for selected task types and behavioral capabilities, excluding discovery sampling. Early, middle, and late stages correspond to training steps 1–50, 51–100, and 101–150, respectively.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Value
Training steps
150
Tasks per step B
32
Rollouts per task G
8
Context length
131,072 tokens
Sampling temperature
1.0
Top- p
1.0
Appendix
Table 2: Training and co-evolution configuration.
Capability
Behavioral focus
Understanding
Interpreting task requirements, constraints, and available information.
Execution
Carrying out intended actions through appropriate tool use and artifact manipulation.
Verification
Checking results against requirements and grounding completion claims in observable evidence.
Debugging
Diagnosing failures, identifying their causes, and applying targeted corrections.
Efficiency
Avoiding redundant work and unnecessary resource use while preserving task progress.
Appendix
Table 3: Behavioral capability categories used to group rubric evidence.
Domain
Tasks
Proportion (%)
Software Engineering
543
31.4
Cybersecurity
415
24.0
Data Analytics
294
17.0
Industrial Engineering
95
5.5
Scientific Research
88
5.1
Digital Media
77
4.5
Appendix
Table 4: Training corpus distribution by broad domain.
Task
Required operations
Assigned task types
Bibliography verification
Check citation metadata against authoritative publication records and identify incorrect venues or years.
Evidence conformance assessment
Clinical study assessment
Assess a case-control study against the supplied Newcastle–Ottawa Scale guidelines and report scores with supporting evidence.
Evidence conformance assessment
3D part mass calculation
Isolate the main mesh component, calculate its volume, and convert units before applying the material density to obtain mass.
3D asset analysis; Unit harmonization
Python build repair
Diagnose compatibility failures, apply targeted patches, and rerun the failing tests to verify the repair.
Artifact diagnosis and repair
Deployment command synthesis
Inspect branch configuration and local changes to construct a deployment command without executing it.
Procedure and plan synthesis
Appendix
Table 5: Representative training tasks and their task-type annotations.
Model
Run 1
Run 2
Run 3
Mean ± Std
Proprietary Models
GPT-5.4 Mini
35.6
33.3
34.5
34.5±1.1
Claude Opus-4.7
59.8
57.5
58.6
58.6±1.1
Open-Weight Models
MiniMax-M2.7
28.7
26.4
31.0
28.7±2.3
GLM-5.1
52.9
50.6
57.5
53.6±3.5
Appendix
Table 6: SkillsBench v1.1 pass rates (%) over three evaluation runs, with the mean and sample standard deviation. Bold values indicate the best pass rates within comparable-scale models.
Evidence. In a feature-planning task, the agent read a format contract but generated headings that did not match the required pattern and omitted dependency fields. Related failures appeared in compliance reports with prescribed templates. The initial criterion drew on nine issue examples and four contrasting examples.
Rubric. When the agent has read an exact output specification, fail if the generated artifact deviates from it and no post-write inspection checks conformance. Pass if the artifact matches the specification or the agent reconciles it against the specification through a read-back or validation check.
Paired skill. Before finalizing, read back the artifact and compare its fields, names, types, headers, and formatting against the documented requirements. Correct deviations and recheck. A successful write confirms file creation, not specification conformance; repeat a successful check only after a relevant modification or a newly discovered deviation.
Lifecycle. Step 10: create the pair with the skill hidden. Step 20: activate the skill at a rubric pass rate of 27.7%. Step 30: hide the skill at 34.3%, retaining the rubric for evaluation. Step 100: retain the pair with the skill hidden at 67.4%. Step 150: retain the pair with the skill hidden at 66.0%. The skill text remains unchanged throughout this interval.
Appendix
Table 7: Artifact specification checking (Verification): guidance is activated and later hidden while the pair is retained.
Evidence. In a telemedicine accessibility task, a text-to-speech tool returned test-mode metadata without the requested audio. The agent wrote placeholder files in place of real audio rather than investigating the missing output. Five issue examples and one contrasting example motivated the new pair.
Rubric. When a tool response lacks the expected data payload, fail if the agent creates placeholder or fabricated artifacts without inspecting the response. Pass if the agent recognizes the missing payload and investigates an alternative way to obtain the real output or reports that it cannot be produced.
Paired skill. Inspect the response content before writing an output file. Check for empty responses, test-mode markers, or metadata without the required artifact. If the payload is absent, investigate alternatives, retry with different parameters, or report the limitation. Do not substitute placeholder files or fabricated content for the missing result.
Lifecycle. Step 40: create the pair with the skill hidden. Step 50: activate the skill at 25.7%. Step 60: hide the skill at 43.2%, retaining the rubric for evaluation. Step 100: retire the rubric–skill pair at 97.2%, removing the criterion from active evaluation.
Appendix
Table 8: Tool-response inspection (Verification): guidance is activated and then hidden before the pair is retired.
Evidence. In tooling-migration and production-scheduling tasks, the agent repeatedly submitted invalid arguments to the todowrite tool after receiving schema errors. Rather than using the error messages to correct the argument structure, it retried the same or substantially similar invalid calls.
Rubric. When a tool returns a schema, argument-validation, or format error, fail if the agent retries with identical or substantially similar invalid arguments without inspecting the concrete error. Pass if it inspects the error and corrects the argument structure on the next attempt, or does not retry the failed call.
Paired skill. Read the full error message and identify the invalid argument, such as a missing field, wrong type, or incorrect nesting. Correct the problematic argument before retrying instead of resubmitting the same payload. Use the error details and documented schema to guide the correction.
Lifecycle. Step 80: create the pair with the skill hidden. Step 90: retain the pair at 67.8%. Step 120: retain it at 89.6%, below the retirement threshold. Step 150: retain it at 81.1%. Across subsequent updates, pass rates remain between the activation and retirement thresholds, so the rubric continues to provide feedback while the skill remains hidden and unchanged.
Appendix
Table 9: Targeted correction of tool arguments (Debugging): the rubric remains active while its paired skill stays hidden.
Evidence. In a portfolio-analysis task, the agent repeatedly reread documentation and data without creating the required deliverables. Other trajectories repeated searches without obtaining new information. Three issue examples and two contrasting examples motivated a criterion focused on redundant exploration.
Rubric. After information gathering begins, fail if the agent repeats reads or searches without obtaining new information or making progress on implementation. Pass if it proceeds to implementation or further inspection yields new information, including targeted rereading that resolves a specific uncertainty.
Paired skill. Once the procedure, inputs, and output requirements are known, proceed to implementation. Reuse information from earlier reads instead of repeating broad searches. If a detail remains unclear, reread the relevant section rather than the entire file, and keep track of resources already inspected.
Lifecycle. Step 130: create the pair with the skill hidden. Step 140: retire the pair at 96.5%. The skill was never activated. Although recent failures motivated the criterion, subsequent evaluations met the retirement threshold, so the pair did not remain in the active pool.
Appendix
Table 10: Avoiding redundant exploration (Efficiency): a late-added criterion is retired without activating its paired skill.
Figure 6: Behavioral retention after rubric retirement. From left to right, the rubrics retire at steps 60, 50, and 80. Open diamonds show training-window pass rates at retirement, while filled circles show fixed-task evaluations without Arise -generated skill guidance.
Method
Average time per step (minutes)
Outcome-only RL
16.46
Arise
17.14
Appendix
Table 11: Average end-to-end training time per step under matched hardware and rollout budgets.
Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.
LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates. Existing skill self-evolution methods directly revise skill text based on execution feedback, but each oracle evaluation requires a full agent rollout, creating a supervision bottleneck that confines search to failure-patching updates. Our key insight is that ranking is a smoother supervision target than absolute outcome regression: identifying which skill is better requires fewer oracle evaluations than predicting exact scores. Building on this insight, we propose SkillLift, which decouples skill search from oracle cost by learning an oracle-aligned rubric as a structured evaluation space. We formalize this as a bilevel optimization problem solved via alternating optimization: an inner loop uses the frozen rubric as a cheap surrogate to guide skill revision at no oracle cost, while an outer loop invokes a small number of oracle rollouts to re-align the rubric via rank correlation, amortizing oracle cost and stabilizing text-space updates. Experiments on complex agent task benchmarks show that our method outperforms existing auto-skill methods with 40--70% less token cost compared to frontier evolving methods. Codes are available at https://github.com/WalteR-MittY-pro/SkillLift.
A persistent skill library allows language model agents to reuse successful strategies across tasks. Maintaining such a library requires three coupled capabilities. The agent selects a relevant skill, utilizes it during execution, and distills new skills from experience. Existing methods optimize these capabilities in isolation or with separate reward sources, resulting in partial and conflicting evolution. We propose Skill1, a framework that trains a single policy to co-evolve skill selection, utilization, and distillation toward a shared task-outcome objective. The policy generates a query to search the skill library, re-ranks candidates to select one, solves the task conditioned on it, and distills a new skill from the trajectory. All learning derives from a single task-outcome signal. Its low-frequency trend credits selection and its high-frequency variation credits distillation. Experiments on ALFWorld and WebShop show that Skill1 outperforms prior skill-based and reinforcement learning baselines. Training dynamics confirm the co-evolution of the three capabilities, and ablations show that removing any credit signal degrades the evolution.
Yaorui Shi, Yuxin Chen, Zhengxi Lu +6
University of Science and Technology of China · 2Meituan · 3National University of Singapore +2