RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers
Organizations: Harvard University · Georgia Institute of Technology
Abstract
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineering capabilities. Real-world robotics extends beyond control: agents must build, integrate, diagnose, and improve heterogeneous artifacts under resource constraints and reason from multimodal feedback. To evaluate these broader capabilities, we introduce RLE-Bench, a benchmark of robot-learning tasks spanning four representative robotics development workflows: interactive control, policy learning, perception and estimation, and mechanical design. We use diverse task-specific metrics to evaluate the artifacts submitted by the coding agents, from the success rate the agents achieved to the policy agents trained, the harness agent built, and the mechanical structures the agent designed. We aggregate these metrics into an overall RLE Index and report workflow-specific capability profiles, enabling systematic comparison of coding agents' capabilities across multiple capability dimensions. Beyond performance ranks, we also conduct in-depth case studies examining agent behavior on representative tasks, highlighting both current capabilities and limitations, and pointing to the opportunities robotics tasks have to offer for future agent training.
Figures & tables
| Benchmark | Coding Agents | Physics Grounded | Controller Synthesis | Robot Policy Training | Perception and Estimation | Mechanical Design |
| SWE/MLE benchmarks | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| LIBERO/RoboTwin/RoboDojo | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ |
| CaP-X | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ |
| RLE-Bench | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Task Family | Workflow | Objective | Submitted artifact | Metrics |
| Interactive control | RoboCasa task learning | Agent context | Success rate | |
| Interactive control | Reusable RoboCasa harness | Harness code + manual | Success rate | |
| Interactive control | Tabletop physical reasoning | – | Solution optimality | |
| Policy learning | Whole-body motion tracking | Model checkpoint | Motion tracking error | |
| Policy learning | VLA recipe engineering | Model checkpoint | Validation loss | |
| Perception & estimation | Blind multi-shape pose estimation | Pose estimator code / models | Estimation error |
| Model | Harness |
| Claude Fable 5.1 ( Anthropic, 2026a ) | Claude Code |
| Claude Opus 5 ( Anthropic, 2026c ) | Claude Code |
| Claude Opus 4.8 ( Anthropic, 2026b ) | Claude Code |
| GPT-6 Astra ( OpenAI, 2026b ) | Codex CLI |
| GPT-5.6 Sol ( OpenAI, 2026a ) | Codex CLI |
| GPT-5.6 Terra ( OpenAI, 2026a ) | Codex CLI |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| System | MPJPE (mm) | Survival (%) | Clearance (%) | Parts/min |
| GPT-6 Astra | 32.62 | 98.71 | 100.00 | 9.22 |
| Claude Fable 5.1 | 40.42 | 100.00 | 86.90 | 14.58 |
| Claude Opus 5 | 34.90 | 93.39 | 78.69 | 6.02 |
| GPT-5.6 Sol | 51.50 | 87.03 | 96.69 | 8.35 |
| DeepSeek-V4.1-Flash | 43.52 | 91.41 | 87.09 | 6.82 |
| Grok 4.6 | 62.22 | 44.77 | 71.52 | 5.89 |
| System | Mean $ | Total $ | Coverage | Mean h | Input M | Cached M | Output M | Failed |
| GPT-6 Astra | 39.00 | 3899.09 | 51/51 | 2.42 | 2601.41 | 2552.76 | 12.07 | 0 |
| Claude Fable 5.1 | 22.81 | 2246.63 | 51/51 | 2.50 | 2671.29 | 2630.17 | 22.51 | 0 |
| Claude Opus 5 | 27.01 | 3143.92 | 51/51 | 2.53 | 5026.40 | 4968.04 | 27.91 | 0 |
| GPT-5.6 Sol | 17.17 | 1535.96 | 51/51 | 2.26 | 2872.33 | 2821.53 | 8.59 | 0 |
| DeepSeek-V4.1-Flash | 0.40 | 39.34 | 51/51 | 2.31 | 3372.05 | 3357.09 | 23.19 | 0 |
| GLM-5.3 Flash | 1.16 | 118.52 | 51/51 | 4.93 | 2994.11 | 2891.37 | 32.74 | 7 |