Cross-Benchmark Transfer from RL on Agentic Coding Tasks
Organizations: Surge AI
Abstract
Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic coding tasks closes this gap, and whether what the agent learns transfers beyond the training distribution. We post-train Kimi K2.7 Code, a 1T-parameter (32B active) open-weight mixture-of-experts model, with RL alone on 1,700 tasks: 1,000 repository tasks graded by hidden fail-to-pass tests and by pass-to-pass tests of existing behavior, and 700 terminal tasks graded by expert-written hidden verifiers. The reward is the fraction of target checks passed and drops to zero if any pass-to-pass test fails. One epoch of GSPO on a rank-32 LoRA adapter improves pass@1 on each of the six external benchmarks we evaluated, across three agent harnesses: SWE-Bench Pro (60.1 to 64.8), DeepSWE (31.0 to 43.4), Terminal-Bench 2.1 (67.4 to 82.0), Terminal-Bench 3 (1.4 to 12.1), Terminal-Bench 4 (0.0 to 7.6), and SWE-Marathon (5.0 to 25.0). Pooled over the five independent task sets (Terminal-Bench 4 revises Terminal-Bench 3), the improvement is significant (p < 0.001), and it remains significant on the three sets released after the training data was collected (p = 0.004); the model also improves under both harnesses never used in training. Median trajectories on DeepSWE and Terminal-Bench 3 are 24-35% shorter in agent steps. The base model's failed DeepSWE runs are mostly near-misses, and on the tasks the trained model newly solves, paired trajectories show it avoiding each of the four failure modes above.
Figures & tables
| Benchmark | Tasks | What it asks for | Harness | Departure from training |
|---|---|---|---|---|
| SWE-Bench Pro | 731 | Issue resolution in 11 repositories (public split); human-verified problems that require multi-file changes | mini-swe-agent | Same format |
| DeepSWE v1.1 ∗ | 113 | Original features in 91 repositories and five languages; reference solutions average 668 added lines | mini-swe-agent | Larger features |
| Terminal-Bench 2.1 | 89 | Hard command-line tasks in software engineering, system administration, scientific computing, security, and data science | Terminus 2 | Same format; unseen harness |
| Terminal-Bench 3 ∗ | 74 | Harder, longer tasks in seven domains (software, science, ML, operations, security, hardware, media); open internet | mini-swe-agent | Harder tasks; open internet |
| Terminal-Bench 4 ∗ | 66 | Terminal-Bench 3 with 8 tasks removed and 19 fixed; 8-hour agent limit | mini-swe-agent | As Terminal-Bench 3 |
| SWE-Marathon v1.1 ∗ | 20 | Complete systems (library reproductions, product clones, ML engineering, algorithmic optimization); agent limits up to 10 hours; expert estimates up to 400 hours | Claude Code | Multi-hour horizon; unseen harness |
| Benchmark | Harness | Tasks | Base | +RL | Gain (pp) |
|---|---|---|---|---|---|
| SWE-Bench Pro | mini-swe-agent | 731 | 60.1 | 64.8 | +4.7 |
| DeepSWE | mini-swe-agent | 113 | 31.0 † | 43.4 | +12.4 |
| Terminal-Bench 2.1 | Terminus 2 | 89 | 67.4 † | 82.0 | +14.6 |
| Terminal-Bench 3 | mini-swe-agent | 74 | 1.4 | 12.1 ‡ | +10.7 |
| Terminal-Bench 4 | mini-swe-agent | 66 | 0.0 | 7.6 | +7.6 |
| SWE-Marathon | Claude Code | 20 | 5.0 | 25.0 | +20.0 |
| Stage | Failure mode before RL | Behavior after RL | Property |
|---|---|---|---|
| Understand | Lost requirements. Built most of the feature but dropped a required constraint, often a subtle one. | Implements the entire specification. Carries every stated requirement through implementation and final review. | (P1) Gradable requirements |
| Test | Narrow testing. Tested the cases its own incomplete implementation was built around. | Tests the requirement, not the implementation. Derives positive, negative, and alternate-form cases from the specification. | (P2) Hidden graders |
| Maintain | Silent regressions. Satisfied the new requirement while breaking behavior that had to stay intact. | Protects existing behavior. Tracks what must change and what must keep working before editing. | (P3) Protected behavior |
| Verify | Weak ground truth. Validated against a false assumption or a weak proxy. | Reconstructs ground truth. Builds an independent reference, an emulator, or its own test inputs. | (P4) No oracle |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Base | +RL | 95% CI of gain (pp) | |
|---|---|---|---|---|
| SWE-Bench Pro | 439/731 | 474/731 | [ 0.2, 9.7] | 0.059 |
| DeepSWE, public baseline | 35/113 | 49/113 | [ 0.2, 24.4] | 0.055 |
| DeepSWE, in-house baseline | 30/113 | 49/113 | [4.4, 28.5] | 0.008 |
| Terminal-Bench 2.1 | 60/89 | 73/89 | [1.8, 26.8] | 0.028 |
| Terminal-Bench 3 | 1/74 | 8/66 | [2.5, 20.8] | 0.010 |
| Terminal-Bench 4 | 0/66 | 5/66 | [0.6, 16.5] | 0.029 |