DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents
Organizations: Nanyang Technological University · Nanjing University
Abstract
Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised. We propose DynaHarness, a dynamic physical harness that couples semantic reasoning with physical governance through a shared execution contract and turns failure evidence into validated capability revisions. To be more specific, the slow brain proposes capabilities and symbolic arguments, while the fast brain grounds and monitors commands, refuses unresolved actions, substitutes capabilities, and requests replans when needed. The physical execution contract bounds each accepted command and records execution evidence across analytic skills, recovery skills, and the frozen VLA. Failure attribution localizes faults in these records and directs targeted revisions of reusable capabilities or execution mechanisms. Paired regression checks govern admission or rejection, closing the self-evolution loop. On LIBERO-Pro, DynaHarness achieves 75.2% on 800 newly sampled initial states, compared with 17.5% for the frozen policy. With the same capability library, full dynamic execution reaches 74.0% versus 63.9% under nominal one-step replanning. This demonstrates the value of DynaHarness as a dynamic physical harness that governs how existing capabilities are grounded, monitored, and coordinated during execution. Our project page is at https://denghaoyuan123.github.io/Dynaharness_page/.
Figures & tables
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
| Command | Steps | Note |
|---|---|---|
| Approach | 30 | |
| Guarded descent | about 30 | 41 before the speed work |
| Close the jaw | 10 | s gripper command; 20 at s |
| Lift | 23 | |
| Carry over the corridor | 60–80 | depends on the corridor height |
| Lower | 15 |
| System | Upper model | Low-level policy | Evaluation setting |
|---|---|---|---|
| DynaHarness | Qwen3-VL-4B | frozen | four cells, episodes each |
| (our run) | none | frozen | four cells, episodes each |
| PhyAgentOS | Qwen3-VL-4B; GPT-4o-mini | frozen | four cells, episodes each |
| Harness VLA | Qwen3-VL-4B | frozen | four cells, episodes each |
| ENPIRE | Qwen3-VL-4B | via generated code | four cells, episodes each |
| Zetta | Qwen3-VL-4B | via tool interface | four cells, episodes each |
| Row | Source | Location and note |
| OpenVLA, | Zhou et al. (2025) | Tables 2 and 4, Average rows (Task and Pos columns); 50 episodes per task |
| (our run) | ours | frozen LIBERO checkpoint, seeds – , all episode rows retained |
| -SFT | Zhang et al. (2026b) | Table 3, row (RLinf pi05_libero130_fullshot ) |
| Fast-WAM | authors; Yuan et al. (2026) | LIBERO-Pro cells supplied by the authors; LIBERO from Table 2 of the publication |
| CaP-Agent0 | Fu et al. (2026) ; Lu et al. (2026) | Goal cells: CaP-X Table 7 Average ( , ); 10-T and 10-S: ASPIRE Table 5 ( , ) |
| RHO | Elmaaroufi et al. (2026) | Table 7 Average, RHO columns ( , ) |
| Reasoner | Goal-T | Goal-S | Reasoner | Goal-T | Goal-S |
|---|---|---|---|---|---|
| GPT-5.5 (low) | 28 | 42 | Gemini 3.1 Pro | 22 | 44 |
| GPT-5.5 (medium) | 24 | 44 | Claude Haiku 4.5 | 20 | 38 |
| GPT-5.5 (high) | 26 | 38 | Claude Sonnet 4.6 | 30 | 42 |
| Gemini Rob-ER 1.6 | 34 | 44 | Claude Opus 4.7 | 22 | 44 |
| Gemini 3.5 Flash | 28 | 48 |
| Method | Upper VLM | Spatial | Object | Goal | Long | Avg. |
| Backbone policies | ||||||
| OpenVLA | none | |||||
| none | ||||||
| (our run) | none | 99.0 | 98.6 | 98.00 | ||
| -SFT | none | 99.0 | ||||
| , task-adapted | none | 99.0 | 98.4 | |||
| Outcome | Episodes |
|---|---|
| Success (benchmark verdict true) | 541 |
| Policy owned the episode | 166 |
| Infrastructure (socket dropped) | 88 |
| Plan refused | 4 |
| Other harness failure | 1 |
| Total, the reported denominator | 800 |
| Zetta round | Goal-T | Goal-S | Mean |
| 31.0 | 38.0 | 34.5 | |
| 1 | 67.5 | 39.5 | 53.5 |
| 2 | 89.5 | 70.5 | 80.0 |
| 3 | 92.0 | 83.0 | 87.5 |
| 4 | 92.5 | 89.0 | 90.8 |
| DynaHarness round | Episodes | Success (%) | |
| Method | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| 99.0 | 99.0 | 98.6 | 95.4 | 98.00 | |
| LIBERO stdlive8c | 99.0 | 99.2 | 96.4 | 93.6 | 97.05 |
| LIBERO formal1 | 98.4 | 99.6 | 97.4 | 94.8 | 97.55 |
| LIBERO formal2 | 98.8 | 99.6 | 97.0 | 96.8 | 98.05 |
| Frozen champion ( 088ef2ea ) | 86.2 | 30.4 | 69.8 | 0.4 | 46.70 |
| Method | Seeds | Goal-T | Goal-S | 10-T | 10-S | Avg. |
|---|---|---|---|---|---|---|
| , frozen | – | 20.5 | 20.5 | 21.0 | 6.5 | 17.1 |
| DynaHarness without evolution | – | 10.5 | 20.0 | 20.0 | 5.0 | 13.9 |
| , frozen | – | 19.5 | 17.5 | 20.0 | 8.0 | 16.2 |
| DynaHarness , analyzed round | – | 75.0 | 69.5 | 73.0 | 53.0 | 67.6 |
| DynaHarness , stat39 | – | 74.5 | 76.0 | 76.0 | 64.5 | 72.8 |
| DynaHarness , archived final | – | 75.0 | 81.0 | 76.0 | 65.0 | 74.25 |
| Perturbation category | Successes | Success (%) |
| Camera viewpoints | 1063/1599 | 66.5 [64.1, 68.8] |
| Robot initial states | 1180/1550 | 76.1 [73.9, 78.2] |
| Objects layout | 1322/1525 | 86.7 [84.9, 88.3] |
| Sensor noise | 1391/1601 | 86.9 [85.1, 88.4] |
| Language instructions | 1360/1537 | 88.5 [86.8, 90.0] |
| Light conditions | 1098/1142 | 96.1 [94.9, 97.1] |
| Suite | C: DynaHarness | C: | B: DynaHarness | B: |
|---|---|---|---|---|
| Goal-T | 79.0 | 26.5 | 78.1 | 26.0 |
| Goal-S | 78.0 | 19.5 | 81.8 | 25.8 |
| 10-T | 77.5 | 17.5 | 75.3 | 7.8 |
| 10-S | 66.5 | 6.5 | 71.1 | 2.2 |
| Episodes | 800 | 800 | 261 | 261 |
| Success (%) | 75.2 | 17.5 | 77.0 | 16.5 |
| Block | Comparison | W/L | (pp) | CI | ||
|---|---|---|---|---|---|---|
| C | Champion vs. | 800 | 473/11 | 57.75 | [45.75, 70.00] | |
| B | Champion vs. | 261 | 161/3 | 60.54 | [44.44, 75.27] | |
| C | Champion vs. stat23 | 800 | 137/22 | 14.38 | [6.62, 23.62] | |
| C | Champion vs. stat28 | 800 | 93/18 | 9.38 | [3.88, 16.00] | |
| C | Champion vs. stat39 | 800 | 12/4 | 1.00 | [0.00, 2.62] | 0.0768 |
| C | stat28 vs. stat23 | 800 | 65/25 | 5.00 | [0.12, 11.50] |
| Arm | Goal-T | Goal-S | 10-T | 10-S | Total | Flags |
|---|---|---|---|---|---|---|
| Full DynaHarness | 150 | 161 | 152 | 129 | 592 | 7 |
| Nominal replanning (A2static) | 130 | 150 | 117 | 114 | 511 | — |
| Frozen sequence (A2seq) | 130 | 159 | 115 | 106 | 510 | 0 |
| No analytic contact skills | 54 | 32 | 31 | 16 | 133 | 8 |
| No recovery/intervention | 149 | 161 | 145 | 130 | 585 | 7 |
| Neither capability group | 47 | 31 | 36 | 12 | 126 | 4 |
| Comparison (Full minus treatment unless named) | W/L | (pp) | CI | Exact |
|---|---|---|---|---|
| Nominal replanning (A2static) | 89/8 | 10.125 | [3.375, 18.250] | |
| Frozen sequence (A2seq) | 93/11 | 10.25 | [3.50, 18.25] | |
| A2static minus A2seq | 11/10 | 0.125 | [ , 2.875] | 1.00 |
| No analytic contact skills | 469/10 | 57.38 | [45.62, 69.00] | |
| No recovery/intervention | 15/8 | 0.88 | [ , 2.62] | 0.21 |
| Neither capability group | 473/7 | 58.25 | [46.62, 69.62] |
| A. Prior routing archive | ||||
|---|---|---|---|---|
| Arm/block | Event stores | Zero VLA (%) | VLA calls/ep. | Planner calls/ep. |
| A2 full control | 770 | 74 | 2.39 | 3.68 |
| No analytic contact skills | 800 | 0 | 18.34 | 4.12 |
| No recovery/intervention | 750 | 75 | 2.11 | 4.09 |
| Neither capability group | 800 | 0 | 18.52 | 4.06 |
| VLA unavailable | 705 | 0.00 | 17.01 | |
| Metric | Full | A2seq |
|---|---|---|
| A. Whole run: episodes per arm | ||
| Planner calls | 2,875 | 800 |
| Environment steps | 244,523 | 188,038 |
| Environment-budget exhaustion | 201 | 17 |
| Tick-budget exhaustion | 7 | 0 |
| B. Common event traces: pairs | ||
| First divergence | Pairs | F | S | Both succeed | Both fail |
|---|---|---|---|---|---|
| Identical dispatches | 468 | 0 | 0 | 468 | 0 |
| Spent-capability substitution | 109 | 16 | 3 | 12 | 78 |
| Planned recovery insertion | 84 | 18 | 1 | 0 | 65 |
| Initial planner choice differs | 61 | 22 | 6 | 19 | 14 |
| Scheduler recovery insertion | 27 | 2 | 0 | 0 | 25 |
| Online planner selection | 19 | 10 | 1 | 0 | 8 |
| Component | Calls | p50 | p90 | p99 | max |
|---|---|---|---|---|---|
| Slow brain (Qwen3-VL-4B) | 3,090 | 913 ms | 1,277 ms | 1,865 ms | 5,581 ms |
| Fast brain | 47,199 | 0.046 ms | 0.056 ms | 0.071 ms | 5.58 ms |
| Fast brain, learned artifact | – | 0.49 ms | – | – | – |
| Episode wall clock | Mean | p50 | p90 | Success | Failure |
| Goal-T | 19.2 s | 15.5 s | 36.0 s | 13.7 s | 35.4 s |
| Goal-S | 22.1 s | 18.5 s | 35.0 s | 18.8 s | 36.6 s |
| Variant | Success (%) | Net episodes | Paired |
|---|---|---|---|
| Latching control | 73.1 | — | — |
| Without latched verdict | 69.6 | ||
| Five-capability control | 73.6 | — | — |
| Without five evolved capabilities | 64.8 |
| Qwen3-VL-4B | Claude Sonnet 5 | Ratio | |
| Latency per call, median (ms) | 1,124 | 6,683 | 5.9 |
| Latency per call, maximum (ms) | 1,485 | 10,348 | 7.0 |
| Prompt tokens per call | 4,036 | 74,600 | 18 |
| Prompt tokens per episode | 21,253 | 444,102 | 20.9 |
| Slow-brain calls per episode | 5.35 | 5.35 | 1.0 |
| Calls over the s timeout | 0 | 0 | – |
| Capability | Inv. | Compl. | Cells | Conversion |
| push_object | 1 | 0 % | 1 | goal_swap[5] |
| slide_drawer_object | 21 | 0 % | 2 | 10_task[3] |
| turn_knob_object | 33 | 21 % | 1 | 10_task[2] |
| keyframe_recovery | 35 | 100 % | 2 | cross-cell |
| swing_door_object | 1 | 0 % | 1 | 10_swap[9] |
| † later rounds, for which only aggregates were retained. | ||||
| Cell | Agent | First change | Both | Paired outcome |
|---|---|---|---|---|
| goal_swap[0] (reachable) | 11 | 18 | 20 | 2 pairs won, none lost |
| goal_swap[3] (unreachable) | 2 | 3 | 3 | 2 against 2 |
| goal_task[7] (unreachable) | 18 | 17 | 19 | within repeat noise |
| 17 further cells (unreachable) | 270 | 270 | 270 | no discordant pair |
| Goal-T and Goal-S ( ) | 301 | 308 | 312 |
| Cell | Instruction | Prom. | Rej. | |
| goal_task[1] | plate on stove | 20 | 0 | |
| 10_task[2] | stove on + pan | 19 | 1 | |
| goal_swap[1] | bowl on stove | 17 | 0 | |
| 10_task[8] | moka pot on stove | 17 | 0 | |
| 10_swap[2] | stove on + moka | 7 | 0 | |
| 10_swap[9] | mug in microwave | 0 | 5 |
| Cell | Method | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Goal-T | ( Ding et al., 2026 ) | 0 | 95 | 10 | 0 | 100 | 0 | 20 | 80 | 5 | 0 | 31.0 |
| Zetta ( Ding et al., 2026 ) | 80 | 100 | 95 | 80 | 100 | 100 | 95 | 95 | 80 | 100 | 92.5 | |
| seeds 1–20 | 0 | 15 | 30 | 0 | 25 | 0 | 20 | 80 | 30 | 5 | 20.5 | |
| seeds 21–40 | 0 | 5 | 25 | 0 | 10 | 0 | 15 | 85 | 55 | 0 | 19.5 | |
| DynaHarness 21–40 | 0 | 100 | 100 | 0 | 90 | 95 | 95 | 90 | 95 | 80 | 74.5 | |
| Goal-S | ( Ding et al., 2026 ) | 0 | 60 | 0 | 45 | 0 | 0 | 0 | 100 | 100 | 75 | 38.0 |
| GPT-4o-mini | Qwen3-VL-4B | |||||
|---|---|---|---|---|---|---|
| Suite or cell | F | Fi | V | F | Fi | V |
| Spatial | 100 | 100 | 100 | 98 | 98 | 99 |
| Object | 99 | 100 | 100 | 99 | 99 | 99 |
| Goal | 98 | 98 | 98 | 98 | 98 | 99 |
| Long | 80 | 92 | 95 | 81 | 96 | 97 |
| LIBERO ( ) | 377 | 390 | 393 | 376 | 391 | 394 |
| LIBERO | LIBERO-Pro | |||
| GPT | Qwen | GPT | Qwen | |
| Attempts per episode, mean | 1.07 | 1.04 | 1.30 | 1.16 |
| Episodes retried | 21 | 16 | 67 | 46 |
| Verifier calls | 39 | 25 | 200 | 200 |
| Environment steps per episode, mean | 168 | 158 | 501 | 442 |
| Policy calls per episode, mean | 34.0 | 32.0 | 100.2 | 88.4 |
| Stage | Recorded outcome |
|---|---|
| Online recovery agent | 3 invocations; image read requested as text, 0 tool calls; failed closed with no environment write |
| Offline diagnosis | 3 visual-evidence claims citing images that were not delivered; diagnosis rejected by the validator |
| Recovery execution | proposal accepted at confidence 0.95; reported completed after 78 steps; joint target unmet; episode failed |
| Candidate proposal | 8 of 9 proposals failed validation |
| Same-seed gate | 3 gates, 0 successes in either arm; none passed |
| GPT-5.6-sol candidate | 13/20 against the bare policy’s 6/20; discordant pairs 8 and 1, exact sign test |
| Task, seed | Outcome | Turns | Tools | Policy | Tokens | Time (s) |
|---|---|---|---|---|---|---|
| 1, 1 | success | 31 | 62 | 3 | 3.32 | 348 |
| 1, 3 | failure | 37 | 70 | 4 | 4.51 | 505 |
| 0, 5 | failure | 33 | 81 | 4 | 3.27 | 448 |
| Paired mechanism benchmark | Incumbent | C1 | C2 |
| goal_task[9] | 16 | 16 | 18 |
| 10_task[0] | 17 | 17 | 19 |
| goal_swap[1] | 16 | 18 | 18 |
| 10_swap[8] | 8 | 9 | 9 |
| goal_swap[5] | 10 | 10 | 10 |
| 10_swap[3] | 0 | 0 | 0 |