Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
Organizations: Nanyang Technological University · Ropedia
Abstract
Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time. We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy. If this state can be represented accurately, the decision making can be written entirely in code. We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it. The same code applies across episodes, and different tasks share one library without a VLM or VLA in the loop. Compared with VLAs and Agent Harnesses, we analyze three advantages of COAP: (i) Explicit State: the state can be stored in code; (ii) Execution: code makes decision making controllable, recovers from failures flexibly, and runs fast and cheaply online; (iii) Extensibility: new tasks reuse, inherit, or extend the shared library, so capabilities can accumulate over tasks. These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable. On RoboDojo's 42 bimanual tasks, the resulting library reaches a success rate of 70.24% without a model at test time. The upper bound of COAP lies in how accurately the state is represented for decision making and how robust the code logic is. We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.
Figures & tables
| Agent-as-Policy | Harness VLA | VLA / WAM | COAP (ours) | |
| Explicit State | ||||
| Inspectable | reasoning text | memory text | ✗ black box | ✓ state and rules |
| Persistent | model context | text memory | ✗ recent frames | ✓ code variables |
| Accurate | pixels, some tools | pixels, some tools | ✗ image pixels | ✓ measured values |
| Execution | ||||
| Controllable | steered by prompt | steered by prompt | ✗ weights decide | ✓ code decides |
| Method | Memory | Precision | Open | Generalization | Long-Horizon | Overall | |
| (6) | (8) | (8) | (12) | (8) | SR | Score | |
| Agent Harness | |||||||
| PhysicalRSI [ 15 ] (agent + VLA) | 46.6 | 32.5 | 24.9 | 15.6 | 37.3 | 31.38 | 36.27 |
| GPT-6 Astra [ 56 ] | 38.7 | 4.0 | 31.0 | 30.5 | 8.2 | 22.48 | 28.97 |
| World Action Model (WAM) | |||||||
| Awomo-0.5 [ 3 ] | 41.2 | 33.5 | 22.9 | 23.8 | 26.8 | 29.64 | 35.34 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Outcome | Share | Failed step (reattempted) | Share | Saved |
| Reattempted, succeeded | 12% | grasp missed or slipped | 49% | 17% |
| Reattempted, failed | 22% | dropped while carrying | 32% | 48% |
| Noticed and reported | 42% | stuck on release | 18% | 23% |
| no further method | 37% | press did not latch | 8% | 100% |
| step budget | 4% | pour incomplete | 1% | 100% |
| Hit the step limit | 1% |
| Step limit | Tasks | COAP | Agent Harness | VLA | GPT-as-policy |
| 200–500 | 12 | 78.1 | 47.3 | 21.9 | 26.5 |
| 550–800 | 15 | 78.4 | 30.6 | 38.0 | 26.9 |
| 900–1900 | 15 | 52.0 | 13.2 | 22.1 | 14.8 |
| Method | Approach | Contact | Examples | Tasks |
| Top | from above | across the body | blocks, digits, a lying pen, eggs | 28 |
| Side | level, toward the axis | around the body or neck | a standing bottle, a cup | 6 |
| Handle | toward the handle | around the bar | a basket’s arch, a broom, a mallet | 3 |
| Rim | down onto the wall | one jaw inside, one outside | a bowl, a mug, a plate | 3 |
| Edge | along the thin side | on the two faces | a coin, a bread slice, a key | 4 |
| Cloth pinch | down onto the hem | fabric layers | a long-sleeved top | 1 |
| make_toast | hang_mugs | sweep_blocks | fill_pen_holder | |
| First evaluation | 3.2 h | 3.2 h | 3.9 h | 0.2 h |
| First success | 4.7 h | 5.2 h | 6.4 h | 4.0 h |
| First merge | 5.8 h | 6.8 h | 7.9 h | 2.7 h |
| Success at merges | 50 77 | 7 50 57 | 0 60 50 | 0 20 27 |
| Layer | Holds | Lines | Used by |
| Task programs | order, parameters, special cases | 23.5k | 2.5% |
| Object families | what an object is, how to hold it | 28.5k | 18% |
| Manipulation | pick, carry, place, push; grasping | 5.4k | 78% |
| Perception | look, measure, confirm | 5.3k | 100% |
| Motion | inverse kinematics, planning, collision | 3.3k | 100% |
| Class | Typical cases |
| Perception error | unrecognized randomized look; wrong final check; split or merged object |
| Contact error | slip or drop in the grip; push stops short; tip at release |
| Incomplete logic | no grasp that also reaches the goal; untried order; unchecked collision |
| Poor code design | fixed size filter; fixed grip-width band; fixed match threshold |
| Unknown | no lead in the trace |
| Task | SR | Score | Task | SR | Score |
| Memory | Generalization | ||||
| cover_blocks | 100.0 | 100.0 | push_T | 97.3 | 97.3 |
| press_by_number | 100.0 | 100.0 | stack_blocks | 94.7 | 95.2 |
| swap_T | 100.0 | 100.0 | stack_bowls | 93.3 | 94.0 |
| swap_blocks | 100.0 | 100.0 | fold_clothes | 90.0 | 91.6 |
| imitate_sorting_sequence | 74.7 | 78.5 | pour_liquid_into_cup | 89.3 | 89.3 |