Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence
Abstract
Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirements, conditions, progress, and failure recovery are implicitly encoded in action sequences, making them difficult to inspect or revise. Digital coding agents offer a precedent: LLMs call tools, verify results, and revise from feedback as executable code. The same working pattern of explicit state, manageable execution, and revisable procedures underlies generalization and long-horizon execution in the physical world, letting physical experience return as reusable programs, memory, or evidence. We propose Physical Coding, representing task state and execution as code. Code as World records objects, relations, constraints, and progress; Code as Policy organizes planning, verification, recovery, and execution. We build HexaAnything, which calls perception, planning, and control tools, including VLA/WAM policies, and makes in-the-loop decisions from external feedback. Verified traces become data and memory, enabling evolution from tools and Harness to model weights, architectures, and ultimately hardware and task design. On RoboCasa365, HexaAnything improves Composite-Unseen and overall success over XR-1 VLA, and its Harness-trained HexaModel beats the base on every split, indicating code traces internalize physical execution. On PhyBench and a dual-arm AgileX robot, the agent autonomously completes physics experiments and most tabletop tasks, often faster than published results. We observe data, model, and tool self-evolution; future work targets weight internalization, autonomous redesign of architectures, languages, representations, and tasks, and deployment in manufacturing and science.
Figures & tables
| Split | XR-1 (native) | HexaAnything | Codex |
|---|---|---|---|
| Atomic-Seen | 78.0% | 80.9% | 81.7% |
| Composite-Seen | 54.8% | 61.5% | 59.8% |
| Composite-Unseen | 34.3% | 38.3% | 34.1% |
| Overall | 56.6% | 61.1 % | 59.5% |
| Task | XR-1 (native) | HexaAnything |
|---|---|---|
| LoadKebabSandwich | 17.0% | 48.0% |
| PortionHotDogs | 22.0% | 37.0% |
| WaffleReheat | 53.0% | 70.0% |
| Stream | Content | Records | Tokens |
|---|---|---|---|
| Agent traces | subgoals, tool calls, verifier outcomes, recovery decisions | 1,010 | 42.4M |
| VQA with CoT | RoboCasa365 state and task-completion judgments; general embodied VQA | 12,075 | 41.7M |
| Scene captions | RoboCasa365 scene descriptions | 6,764 | 2.0M |
| General data | math, code, and instruction following | 24,491 | 92.5M |
| Total | 44,340 | 178.6M |
| Planner | Atomic-Seen | Composite-Seen | Composite-Unseen | Overall |
|---|---|---|---|---|
| XR-1 (native, no Harness) | 78.0% | 54.8% | 34.3% | 56.6% |
| Qwen3.8-27B (base) | 80.8% | 61.0% | 37.3% | 60.5% |
| GPT-5.6-Sol | 80.9% | 61.5% | 38.3% | 61.1% |
| HexaModel v0.1 | 81.0% | 62.3% | 39.5% | 61.7% |
| Revision round | Fold cloth | Pour vase | Press by number |
|---|---|---|---|
| Round 0 | 0.0% | 40.0% | 0.0% |
| Round 1 | 60.0% | 0.0% | 0.0% |
| Round 2 | 80.0% | 100.0% | 100.0% |
| Harness + Model | Hooke’s law | Simple pendulum | Coupled oscillators |
|---|---|---|---|
| error | error | error | |
| HexaAnything + GPT-6-Astra | 2.3% (10/10) | 1.2% (10/10) | 0.8% (10/10) |
| HexaAnything + Opus 5.5 | 1.3% (10/10) | 1.0% (10/10) | 1.7% (10/10) |
| HexaAnything + GPT-5.6-Sol | 4.8% (9/10) | 1.9% (8/10) | 1.2% (10/10) |
| HexaAnything + Qwen3.8-Max | 12.4% (5/10) | 2.4% (3/10) | 3.7% (3/10) |
| Task | Success | Progress | Output tokens | Time (min) | Published reference (same model) |
|---|---|---|---|---|---|
| Unscrew bottle cap | 3/3 | 100% | 1.52k | 5.1 | GPT-Policy R3: 3/3, 17.9 min |
| Tic-tac-toe | 3/3 | 100% | 3.26k | 5.0 | GPT-Policy R8: 3/3, 13.6 min |
| Block into bowl | 3/3 | 100% | 443 | 1.0 | Robocurve: 19/20, 2.1k tokens, 2.5 min |
| Hot-dog serving | 3/3 | 100% | 2.40k | 7.9 | – |
| Pour blocks | 3/3 | 100% | 2.29k | 3.1 | – |
| Toss blocks into bowl | 1/3 | 77.8% | 4.69k | 8.1 | – |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Line of work and representatives | Primary evolving artifact | Persistence and typical fixed boundary | Feedback or evidence | Relation to this work |
|---|---|---|---|---|
| Task, scene, and curriculum generation: GenSim, RoboGen, SceneSmith, CurricuLLM [ 82 , 87 , 63 , 70 ] | Task specification , scene/environment , or curriculum | Generated tasks and scenes persist as training or evaluation assets; the downstream policy, learner, and evaluator are usually external. | Task solvability, policy performance, or curriculum progress. | Covers code as world and environment evolution; it does not by itself revise the executing policy and verifier together. |
| Data and trajectory synthesis: GenSim2, MimicGen, RoboTwin, HumanoidGen [ 30 , 56 , 59 , 36 ] | Data : demonstrations, trajectories, captions, or VQA | Data is retained for later training; the generator and downstream update rule are usually fixed for a run. | Dataset scale or quality and downstream policy success. | Motivates data evolution; our traces link training data to state predicates, recovery, and held-out validation. |
| Reward and objective search: Eureka, DrEureka, REvolve [ 55 , 54 , 27 ] | Reward code, domain randomization, or preference-derived objective | Candidate objectives are iterated within training; the policy and evaluation boundary typically remain specified externally. | Rollout return or success plus automated or human feedback. | Shows objective evolution, but a training reward is not an independent verifier . |
| Evaluation and verifier generation: RoboPlayground, AutoEval, SimFoundry [ 86 , 105 , 67 ] | Task specifications, resets, success predicates, and evaluator | Evaluation artifacts can be reused across policies; policy training is usually held fixed during comparison. | Terminal success, process evidence, reset reliability, or evaluator agreement. | Directly informs and provenance; our verifier is coupled to recovery and regression tests while remaining independent. |
| Programmatic policy synthesis: Code as Policies, RoboCodeX, RoboScript, RoboCoder, CaP-X [ 47 , 58 , 10 , 41 , 21 ] | Executable policy programs and skill compositions | Programs execute per task or are reused as skills; the base model, tool semantics, and evaluator are generally fixed. | Execution success, task completion, and code or skill validity. | Closest precedent for code as policy; our world program makes intermediate state and constraints explicit. |
| Tools, interfaces, and agent runtimes: ReAct, SWE-agent, OpenHands, Model Hardware Standard [ 98 , 95 , 85 , 3 ] | Tool schema, agent-computer/robot interface, or hardware driver | The interface structures observations and actions; model weights and task/evaluation protocols generally remain fixed. | Tool execution, repository tests, or interface-level correctness. | Supports modular boundaries and physical adapters, but interface standardization alone is not self-evolution. |
| Record | Initial displacement | (rad/s) | (rad/s) |
|---|---|---|---|
| 1 | CH1 mm | 3.0139 | 3.3382 |
| 2 | CH2 mm | 3.0141 | 3.3383 |
| Estimate (mean of records) | 3.0140 | 3.3382 | |
| Reference | 3.0137 | 3.3376 |