RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation
Authors: Shifeng Bao, Fanding Huang, Yihan Lin, Youhe Feng, Guanlin Li, Chen Zhao, Yang Li, Jiawei He, +2 more
Organizations: School of Information, Renmin University of China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China · Tsinghua University · XYZ Embodied AI, Beijing, China · Engineering Research Center of Database and Business Intelligence, Beijing, China
Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.
Figures & tables
Figure 1: Self-improving robotic manipulation through hierarchical knowledge evolution. (a) In ordered block manipulation, the robot may switch to the next block before completing the current subtask (top), or select the correct target but release the cover off target (bottom). (b) Physical feedback updates Task and Action Knowledge across episodes, with the VLM and executor held fixed. (c) Knowledge evolution improves held-out simulation success by 18.3–26.7 percentage points after 80 rollouts. On real-world tasks, simulation-derived knowledge achieves 41.7% mean zero-shot success, which rises to 62.5% after further knowledge updates from real-world interaction.
Figure 2: Overview of RoboHarn-Evo. Top: The inner loop executes manipulation operations and records their outcomes. Bottom: The outer loop organizes this experience into Task and Action Knowledge for subsequent decisions. Concrete motions are computed from the current scene.
Figure 3: Skill-routed retrieval and evidence-driven maintenance. Top: Task retrieval determines what to do ; conditioned on that subtask, Action retrieval determines how to act using supported atomic knowledge. Bottom: Before–after execution evidence is reflected into new Task and Action Knowledge, assigned to Skills, and used to consolidate and update only the affected knowledge groups.
Task
π0.5
X-VLA
Mem-0
HarnessVLA
Qwen3.8-27B
GPT-5.5
GPT-6
w/o HPK
Full HPK
w/o HPK
Full HPK
w/o HPK
Full HPK
Rearrange Blocks
13%
13%
89%
30%
15%
35%
50%
80%
75%
90%
Swap Blocks
24%
16%
67%
35%
10%
25%
45%
70%
70%
85%
Press Button
0%
0%
0%
75%
35%
60%
95%
100%
100%
100%
Swap T
15%
3%
14%
35%
20%
40%
55%
90%
45%
95%
Put Back Block
11%
18%
90%
25%
20%
50%
50%
75%
75%
90%
Table 1: Task success on RMBench. We report task success rates (%) on six RMBench tasks, with Overall denoting the six-task mean. X-VLA and Mem-0 results are reported from the Mem-0 benchmark evaluation; all remaining results are evaluated on 20 held-out episodes per task. Within each model, w/o HPK and Full HPK share the same executor. HarnessVLA is tested with GPT-5.5.
Figure 4: Knowledge maintenance and held-out manipulation performance. (a,b) Active-knowledge correctness under shared-evidence replay. (c,d) Outcomes of the 12 initial errors and retention of the 24 initially correct entries at K=80 . (e) Held-out success across learning checkpoints. (f) Endpoint comparisons using the same source experience at K=80 . K counts interaction rollouts. Error bars in (e,f) indicate standard deviations over three learning histories. Q2 replays previously collected interaction records to evaluate knowledge maintenance; Q3 collects experience through online interaction to evaluate subsequent task performance.
Task
GPT-5.5
GPT-6
w/o HPK
w/ Full HPK
w/o HPK
w/ Full HPK
Score
SR
Score
SR
Score
SR
Score
SR
Cover Blocks
37.0
30.0
75.0
70.0
57.0
50.0
93.0
90.0
Press by Number
50.0
50.0
80.0
80.0
80.0
80.0
90.0
90.0
Average
43.5
40.0
77.5
75.0
68.5
65.0
91.5
90.0
Δ (pp)
–
+34.0
+35.0
–
+23.0
+25.0
Table 3: Zero-shot knowledge transfer from RMBench to RoboDojo. Score and success rate (SR) are reported as percentages, with 10 target episodes per task. Δ is the difference in Average between w/ Full HPK and w/o HPK , in percentage points.
Figure 5: Real-world transfer and self-improvement. (a) Task sequences and example instructions. (b) Success rate (top) and normalized task progress (bottom) for π0.5 and RoboHarn-Evo before and after RSI. Zero-shot uses simulation-derived HPK; After RSI incorporates additional real-world interaction. RoboHarn-Evo results use eight scenes per task, with full task completion counted as success.
Figure 6: Task-conditioned knowledge reuse from simulation to the real world. (a) Simulated pressing experience provides physical evidence for maintaining HPK. (b) A later real-world episode uses block-derived counts and scene-specific contact geometry to complete the second green press before switching to blue. The right panels show condensed Task and Action Knowledge entries with accumulated source evidence. Arrows illustrate the maintenance and reuse workflow.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Real-world hardware setup. The AgileX Cobot Magic platform uses two follower arms for tabletop manipulation, with one front-view and two wrist-mounted RGB-D cameras. The follower arms are labeled Puppet Left and Puppet Right in the image.
Figure 8: Example of Cover Blocks.
Type
Skill family
Persistent entry
Task
Cover by spatial order
Cover the leftmost visible target
Task
Cover by spatial order
Cover the next (middle) visible target
Task
Cover by spatial order
Cover the remaining rightmost target
Task
Uncover by requested identity
Uncover the red block first
Task
Uncover by requested identity
Uncover the green block next
Task
Uncover by requested identity
Uncover the blue block last
Appendix
Table 4: Complete catalog of the persistent knowledge entries extracted for Cover Blocks . The table summarizes semantic content and skill grouping, rather than treating the independently collected HPK examples as repeated validation trials.
Figure 9: Evidence associated with the two representative entries from demonstration episode_9 . The Task Knowledge evidence spans frames 512–687. The Action Knowledge evidence for grasping the enclosing lid spans frames 512–600. The final frame also verifies completion of the subsequent set-aside action.
Component
Task Knowledge
Action Knowledge
Condition
Red is uncovered; green and blue remain covered.
A lid covers the target block; its top handle is reachable.
Strategy
Uncover green next, according to the requested color order.
Pinch opposite sides of the top handle; lift vertically to clear the block before moving laterally.
Outcome
Green is visible and its lid is moved away; blue remains covered.
The lid rises with the gripper and clears the block.
Evidence
Observations at the subtask boundaries establish the required change in task state.
Observations of the grasp and lift establish control of the lid.
Appendix
Table 7: Two linked knowledge entries for Cover Blocks , with the requested uncovering order red, green, then blue. The grasp effect is a step toward, rather than the completion of, the uncovering subtask.
Category
# Skills
Primary Responsibility
Perception
1
Normalize perception queries before segmentation and scene-memory binding.
Memory
1
Convert observation and execution evidence into compact runtime memory.
Planning
5
Maintain task planning, runtime reasoning, and high-level execution decisions.
Monitoring
1
Detect execution health and semantic OOD states.
Tool Calling
19
Route tool use and construct grounded tool-call procedures from reusable primitives and workflows.
Effect Verification
1
Verify whether executed actions produced their intended physical effects.
Appendix
Table 8: Organization of the runtime skill library.
Skill
Responsibility
Control Turn Planner
Decides the next top-level action forone agent control turn from the task instruction, committed memory, observations, state information, and available runtime state. It returns the next committed memory, an executor-facing subtask, an action mode, an arm preference, and optionally a selected skill.
Runtime State Reasoning
Defines how the planner interprets structured task, working, perception, scene-memory, monitor, manipulation, and recovery state. It treats structured runtime fields as the authoritative state source and specifies how that state should be used when producing the next control decision.
Long Horizon Execution
Handles global tasks that must be decomposed into ordered subtasks. It refines the subtask plan, selects the next subtask, checks the previous subtask, decides whether to retry, recover, continue, or finish, and delegates concrete execution to a monitored execution boundary.
Monitored Subtask Execution
Executes exactly one narrow executor-facing instruction as a monitored action unit. It monitors success, failure, stall, and timeout conditions, performs deterministic stopping or resetting when needed, and returns a structured execution result to the calling workflow.
EAP Data Collection
Implements a forward/reverse EAP-style data-collection workflow. It initializes a collection run, executes forward behavior, executes reverse or reset behavior, keeps the environment reusable, and records structured run and dataset information.
Appendix
Table 9: Planning skills and their responsibilities.
Skill
Responsibility
Tool-Calling Router
Selects one workflow and a post-execution intent from an explicit OOD scenario or monitor signal. It returns a structured routing decision and does not itself execute tools.
Close Gripper
Closes a selected gripper to re-establish or stabilize grasp state after an uncertain, failed, or slipping grasp.
Contact Displace
Applies a very small bounded end-effector displacement when local contact should be released or probed without initiating a new task-level action.
Lift End Effector
Creates bounded vertical end-effector clearance when the arm or gripper is locally blocked or obstructed.
Move EE to Grounded Instance
Moves an end effector toward manipulation geometry associated with a grounded scene-memory instance or a public placement target. Executable pose selection remains runtime-side.
Move EE to Pose
Moves an end effector toward an already validated absolute world-frame target pose or position. It does not perform semantic grounding.
Appendix
Table 10: Routing and primitive skills in the Tool Calling category.
Skill
Responsibility
Go Home and Retry
Handles exhausted or unproductive rollouts by constructing a situation-specific tool sequence that returns the robot toward a reusable posture before control is returned to planning.
Recover Grasp Lost
Constructs a bounded tool sequence when the manipulated object is no longer attached to or controlled by the gripper.
Recover Motion Blocked
Constructs a situation-specific tool sequence when local contact or blockage prevents the current motion from continuing safely.
Recover Object Not Visible
Constructs a tool-use plan that improves observability or returns control to planning when a task-relevant object cannot be reliably observed.
Recover Requires Replan
Returns control to planning when the current subtask is no longer valid. It may produce no physical tool calls when physical intervention is unnecessary.
Recover Scene Drift
Refreshes observation and determines whether execution can be retried or should be replanned when the scene has changed enough to invalidate the current rollout context.
Appendix
Table 11: Workflow skills in the Tool Calling category. Each workflow constructs a situation-dependent sequence from the runtime tools currently available.
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventions through interaction with these policies. We propose Recursive Harness Distillation to accumulate this experience as reusable guidance across agents. A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback. The resulting playbook enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters. In real-world manipulation, the harness improves success from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook. The same playbook also benefits the strong agent, which reaches 79.2% success. These results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.
Seungyeon Kim, Junhoo Lee, Minkyu Kim +2
Seoul National University · Korea Advanced Institute of Science and Technology
Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rules. We refer to this capability as intuitive manipulation. Existing benchmarks fail to capture this integration: they evaluate physical reasoning in isolation from execution, or measure policy performance without requiring explicit reasoning. We introduce IMBENCH, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution. Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies. We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios. Experiments reveal a consistent gap: vision language models show partial physical reasoning ability but fail to produce executable plans, while state-of-the-art vision-language-action models struggle to satisfy task constraints and generalize across scenarios. These results identify intuitive manipulation as a missing axis in current foundation models and generalist robot policies, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.
We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. Across assembly, block construction, and dice flipping, AGP succeeds in at least eight of ten trials for each evaluated task configuration. We further study efficiency through task experience accumulation and find that reusing saved procedures and programs shortens execution time across repeated trials. These findings support a path for general-purpose agents to act as robot policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.
Mengzhao Jia, Yang Lin, Xixin Zhang +3
University of Notre Dame · University of California San Diego · San Diego State University