cs.ROOct 8, 2026

Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement

Authors: Kairui Hu, Siyuan Hu, Fangzhou Hong, Zhaoxi Chen, Ziwei Liu

Organizations: Nanyang Technological University · Ropedia

Abstract

Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time. We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy. If this state can be represented accurately, the decision making can be written entirely in code. We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it. The same code applies across episodes, and different tasks share one library without a VLM or VLA in the loop. Compared with VLAs and Agent Harnesses, we analyze three advantages of COAP: (i) Explicit State: the state can be stored in code; (ii) Execution: code makes decision making controllable, recovers from failures flexibly, and runs fast and cheaply online; (iii) Extensibility: new tasks reuse, inherit, or extend the shared library, so capabilities can accumulate over tasks. These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable. On RoboDojo's 42 bimanual tasks, the resulting library reaches a success rate of 70.24% without a model at test time. The upper bound of COAP lies in how accurately the state is represented for decision making and how robust the code logic is. We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RHO: Your Coding Agent is Secretly a Roboticist

    Jun 15, 2026Karim Elmaaroufi, Justin Svegliato, Sarunas Kalade +3Robot Policy LearningAI Coding Agents

  2. RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning

    Sep 24, 2026Zexi Li, Yehang Zhang, Haojian Huang +10Robot Policy LearningRobot Policy Generalization

  3. Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation

    Jun 2, 2026Nan Sun, Yuan Zhang, Yongkun Yang +10Long-Horizon Robotic ManipulationCoT Reasoning