ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
Organizations: ShanghaiTech University · Ant Group
Abstract
As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric-Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency. The project page is at https://foundation-model-research.github.io/ARISE .
Figures & tables
| SkillsBench v1.1 | TB v2.1 | |||||||||
| Model | SE | NS | OW | IP | FE | MR | CS | MC | Overall | Overall |
| Proprietary Models | ||||||||||
| GPT-5.4 Mini | 27.1 | 35.7 | 45.2 | 21.4 | 25.9 | 41.7 | 33.3 | 66.7 | 34.5 | 59.2 |
| GPT-5.5 | 63.4 | 77.9 | 76.2 | 57.5 | 37.0 | 95.0 | 69.0 | 60.0 | 67.3 | 84.3 |
| Claude Opus-4.7 | 58.3 | 83.3 | 54.8 | 54.8 | 44.4 | 50.0 | 57.1 | 53.3 | 58.6 | 83.1 |
| Open-Weight Models | ||||||||||
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration | Value |
|---|---|
| Training steps | 150 |
| Tasks per step | 32 |
| Rollouts per task | 8 |
| Context length | 131,072 tokens |
| Sampling temperature | 1.0 |
| Top- | 1.0 |
| Capability | Behavioral focus |
|---|---|
| Understanding | Interpreting task requirements, constraints, and available information. |
| Execution | Carrying out intended actions through appropriate tool use and artifact manipulation. |
| Verification | Checking results against requirements and grounding completion claims in observable evidence. |
| Debugging | Diagnosing failures, identifying their causes, and applying targeted corrections. |
| Efficiency | Avoiding redundant work and unnecessary resource use while preserving task progress. |
| Domain | Tasks | Proportion (%) |
|---|---|---|
| Software Engineering | 543 | 31.4 |
| Cybersecurity | 415 | 24.0 |
| Data Analytics | 294 | 17.0 |
| Industrial Engineering | 95 | 5.5 |
| Scientific Research | 88 | 5.1 |
| Digital Media | 77 | 4.5 |
| Task | Required operations | Assigned task types |
|---|---|---|
| Bibliography verification | Check citation metadata against authoritative publication records and identify incorrect venues or years. | Evidence conformance assessment |
| Clinical study assessment | Assess a case-control study against the supplied Newcastle–Ottawa Scale guidelines and report scores with supporting evidence. | Evidence conformance assessment |
| 3D part mass calculation | Isolate the main mesh component, calculate its volume, and convert units before applying the material density to obtain mass. | 3D asset analysis; Unit harmonization |
| Python build repair | Diagnose compatibility failures, apply targeted patches, and rerun the failing tests to verify the repair. | Artifact diagnosis and repair |
| Deployment command synthesis | Inspect branch configuration and local changes to construct a deployment command without executing it. | Procedure and plan synthesis |
| Model | Run 1 | Run 2 | Run 3 | Mean Std |
| Proprietary Models | ||||
| GPT-5.4 Mini | 35.6 | 33.3 | 34.5 | |
| Claude Opus-4.7 | 59.8 | 57.5 | 58.6 | |
| Open-Weight Models | ||||
| MiniMax-M2.7 | 28.7 | 26.4 | 31.0 | |
| GLM-5.1 | 52.9 | 50.6 | 57.5 | |
| Evidence. In a feature-planning task, the agent read a format contract but generated headings that did not match the required pattern and omitted dependency fields. Related failures appeared in compliance reports with prescribed templates. The initial criterion drew on nine issue examples and four contrasting examples. |
| Rubric. When the agent has read an exact output specification, fail if the generated artifact deviates from it and no post-write inspection checks conformance. Pass if the artifact matches the specification or the agent reconciles it against the specification through a read-back or validation check. |
| Paired skill. Before finalizing, read back the artifact and compare its fields, names, types, headers, and formatting against the documented requirements. Correct deviations and recheck. A successful write confirms file creation, not specification conformance; repeat a successful check only after a relevant modification or a newly discovered deviation. |
| Lifecycle. Step 10: create the pair with the skill hidden. Step 20: activate the skill at a rubric pass rate of 27.7%. Step 30: hide the skill at 34.3%, retaining the rubric for evaluation. Step 100: retain the pair with the skill hidden at 67.4%. Step 150: retain the pair with the skill hidden at 66.0%. The skill text remains unchanged throughout this interval. |
| Evidence. In a telemedicine accessibility task, a text-to-speech tool returned test-mode metadata without the requested audio. The agent wrote placeholder files in place of real audio rather than investigating the missing output. Five issue examples and one contrasting example motivated the new pair. |
| Rubric. When a tool response lacks the expected data payload, fail if the agent creates placeholder or fabricated artifacts without inspecting the response. Pass if the agent recognizes the missing payload and investigates an alternative way to obtain the real output or reports that it cannot be produced. |
| Paired skill. Inspect the response content before writing an output file. Check for empty responses, test-mode markers, or metadata without the required artifact. If the payload is absent, investigate alternatives, retry with different parameters, or report the limitation. Do not substitute placeholder files or fabricated content for the missing result. |
| Lifecycle. Step 40: create the pair with the skill hidden. Step 50: activate the skill at 25.7%. Step 60: hide the skill at 43.2%, retaining the rubric for evaluation. Step 100: retire the rubric–skill pair at 97.2%, removing the criterion from active evaluation. |
| Evidence. In tooling-migration and production-scheduling tasks, the agent repeatedly submitted invalid arguments to the todowrite tool after receiving schema errors. Rather than using the error messages to correct the argument structure, it retried the same or substantially similar invalid calls. |
| Rubric. When a tool returns a schema, argument-validation, or format error, fail if the agent retries with identical or substantially similar invalid arguments without inspecting the concrete error. Pass if it inspects the error and corrects the argument structure on the next attempt, or does not retry the failed call. |
| Paired skill. Read the full error message and identify the invalid argument, such as a missing field, wrong type, or incorrect nesting. Correct the problematic argument before retrying instead of resubmitting the same payload. Use the error details and documented schema to guide the correction. |
| Lifecycle. Step 80: create the pair with the skill hidden. Step 90: retain the pair at 67.8%. Step 120: retain it at 89.6%, below the retirement threshold. Step 150: retain it at 81.1%. Across subsequent updates, pass rates remain between the activation and retirement thresholds, so the rubric continues to provide feedback while the skill remains hidden and unchanged. |
| Evidence. In a portfolio-analysis task, the agent repeatedly reread documentation and data without creating the required deliverables. Other trajectories repeated searches without obtaining new information. Three issue examples and two contrasting examples motivated a criterion focused on redundant exploration. |
| Rubric. After information gathering begins, fail if the agent repeats reads or searches without obtaining new information or making progress on implementation. Pass if it proceeds to implementation or further inspection yields new information, including targeted rereading that resolves a specific uncertainty. |
| Paired skill. Once the procedure, inputs, and output requirements are known, proceed to implementation. Reuse information from earlier reads instead of repeating broad searches. If a detail remains unclear, reread the relevant section rather than the entire file, and keep track of resources already inspected. |
| Lifecycle. Step 130: create the pair with the skill hidden. Step 140: retire the pair at 96.5%. The skill was never activated. Although recent failures motivated the criterion, subsequent evaluations met the retirement threshold, so the pair did not remain in the active pool. |
| Method | Average time per step (minutes) |
|---|---|
| Outcome-only RL | 16.46 |
| Arise | 17.14 |