HarnessPAI: An Evolving Harness for Physical AI
Abstract
Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. Yet the field has focused primarily on the last component: the action model that maps observations to low-level controls. The prevailing training recipe can erode the perceptual and reasoning capabilities needed for robust behavior, leaving even strong action models vulnerable to scene perturbations and long-horizon tasks. We introduce HarnessPAI, a model- and embodiment-agnostic Harness framework for Physical AI that treats code as the executable and evolvable interface that organizes the underlying action primitive. The framework separates two timescales: within a rollout, it executes open-loop at the program level, with a fixed program guiding and checking execution; across rollouts, it evolves closed-loop, using execution feedback to revise the program and distill failures into reusable skills. Across desktop robot arms, household robots, a robot vacuum, and a legged walking agent, HarnessPAI improves on both pure action models and code-as-policy baselines without retraining the underlying model: a 61.6-point gain over on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks. Once a program is selected, rollout execution requires no online high-level LLM deliberation. Beyond execution, the converged program is also a cheap and reliable expert-data collector, and fine-tuning on collected expert data lifts success rate on LIBERO-PRO by 38.8 points. Our results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system. Website: https://darwin-agent.github.io/HarnessPAI
Figures & tables
| Method family | Adaptation location | Coding agent during rollout | API-based control | Learned policy usage | Evolution across rollouts |
| VLA/WAM post-training | Model parameter | ✗ | ✗ | ✓ | Training only |
| Code-as-policy | Program | ✗ | ✓ | ✗ | Sometimes |
| VLA harness | Runtime/memory | ✓ | ✓ | ✓ | Usually |
| HarnessPAI | Program/memory | ✗ | ✓ | ✓ | Designed for persistent evolution |
| Model | Role | Responsibility | Edits code |
| GPT-5.6-sol | writer + synthesizer | edits phase code; self-review; next instruction | ✓ |
| Gemini-3.5-Flash | diagnoser | watches rollout video; node-by-node verdict vs. graph | ✗ |
| Phase | Executor |
| grounding | fixed code (segmentation + geometry) |
| approach | fixed code (Cartesian servo, Pyroki IK rescue) |
| grasp | VLA |
| transport | fixed code |
| place | VLA |
| Method | Spatial | Object | Goal | Overall |
| Pure-Model | ||||
| DreamZero-LIBERO ( RLinf, 2026 ) | 81.6 | 90.2 | 41.2 | 71.0 |
| TraceVLA ( Zheng et al., 2025 ) | 84.6 | 85.2 | 75.1 | 81.6 |
| OpenVLA ( Kim et al., 2024 ) | 84.7 | 88.4 | 79.2 | 84.1 |
| WorldVLA ( Cen et al., 2025 ) | 85.6 | 89.0 | 82.6 | 85.7 |
| CoT-VLA ( Zhao et al., 2025 ) | 87.5 | 91.6 | 87.6 | 88.9 |
| Method | Type | Atomic-Seen |
| ( Black et al., 2024 ) | VLA | 34.6 |
| ( Intelligence et al., 2025b ) | VLA | 39.6 |
| RLDX-1 ( RLWRLD, 2026 ) | VLA | 60.0 |
| WorldDreamer ( worldAgents-c, 2026 ) | WAM | 65.0 |
| Harness VLA (Codex) ( Zhang et al., 2026b ) | Harness | 91.6 |
| Harness VLA (CC) ( Zhang et al., 2026b ) | Harness | 79.4 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | 2A-Hand | 2A-Lift | Insert | Lift | Stack | Restack | Wipe | Avg |
| CaP-Agent | 20 | 74 | 0 | 97 | 98 | 89 | 100 | 68.3 |
| Human | 91 | 94 | 80 | 93 | 73 | 47 | 100 | 82.6 |
| ASPIRE | 92 | 71 | 9 | 97 | 99 | 100 | 99 | 81.0 |
| HarnessPAI | 100 | 91 | 83 | 100 | 100 | 100 | 100 | 96.3 |
| vs. ASPIRE | +8.0 | +20.0 | +74.0 | +3.0 | +1.0 | 0.0 | +1.0 | +15.3 |
| Object | Goal | Spatial | |||||
| Method | Swap | Task | Swap | Task | Swap | Task | Avg |
| OpenVLA | 0 | 0 | 0 | 0 | 0 | 0 | 0.0 |
| 0 | 0 | 0 | 0 | 0 | 0 | 0.0 | |
| 29.8 | 10.8 | 37.6 | 29.0 | 51.8 | 50.6 | 34.9 | |
| CaP-Agent | 22 | 18 | 26 | 17 | 12 | 14 | 18.2 |
| ASPIRE | 98 | 95 | 81 | 45 | 51 | 60 | 71.7 |
| Task | Description | Success / Total | Success Rate |
| LIBERO-Object | |||
| Task 0 | Pick up the alphabet soup and place it in the basket | 50/50 | 100.0 |
| Task 1 | Pick up the cream cheese and place it in the basket | 50/50 | 100.0 |
| Task 2 | Pick up the salad dressing and place it in the basket | 50/50 | 100.0 |
| Task 3 | Pick up the BBQ sauce and place it in the basket | 50/50 | 100.0 |
| Task 4 | Pick up the ketchup and place it in the basket | 50/50 | 100.0 |
| Task | Description | Success / Total | Success Rate |
| LIBERO-Object | |||
| Task 0 | Pick up the alphabet soup and place it in the basket | 50/50 | 100.0 |
| Task 1 | Pick up the cream cheese and place it in the basket | 50/50 | 100.0 |
| Task 2 | Pick up the salad dressing and place it in the basket | 50/50 | 100.0 |
| Task 3 | Pick up the BBQ sauce and place it in the basket | 50/50 | 100.0 |
| Task 4 | Pick up the ketchup and place it in the basket | 50/50 | 100.0 |
| Task | Description | Success / Total | Success Rate |
| LIBERO-Object | |||
| Task 0 | Pick up the alphabet soup and place it in the basket | 45/50 | 90.0 |
| Task 1 | Pick up the cream cheese and place it in the basket | 50/50 | 100.0 |
| Task 2 | Pick up the salad dressing and place it in the basket | 50/50 | 100.0 |
| Task 3 | Pick up the BBQ sauce and place it in the basket | 46/50 | 92.0 |
| Task 4 | Pick up the ketchup and place it in the basket | 41/50 | 82.0 |
| Swap | Task | |||||
| Task | Succ./Total | Fail | Rate | Succ./Total | Fail | Rate |
| LIBERO-Object | ||||||
| Task 0 | 50/50 | 0 | 100.0 | 50/50 | 0 | 100.0 |
| Task 1 | 50/50 | 0 | 100.0 | 50/50 | 0 | 100.0 |
| Task 2 | 50/50 | 0 | 100.0 | 47/50 | 3 | 94.0 |
| Task 3 | 50/50 | 0 | 100.0 | 50/50 | 0 | 100.0 |
| Swap | Task | |||||
| Task | Succ./Total | Fail | Rate | Succ./Total | Fail | Rate |
| LIBERO-Object | ||||||
| Task 0 | 41/50 | 9 | 82.0 | 0/50 | 50 | 0.0 |
| Task 1 | 50/50 | 0 | 100.0 | 50/50 | 0 | 100.0 |
| Task 2 | 0/50 | 50 | 0.0 | 0/50 | 50 | 0.0 |
| Task 3 | 5/50 | 45 | 10.0 | 0/50 | 50 | 0.0 |
| Task | Name | Layout / Style | Success / Total | Success Rate |
| Task 0 | CloseBlenderLid | 8 / 8 | 19/20 | 95.0 |
| Task 1 | CloseFridge | 2 / 2 | 20/20 | 100.0 |
| Task 2 | CloseToasterOvenDoor | 7 / 7 | 19/20 | 95.0 |
| Task 3 | CoffeeSetupMug | 4 / 4 | 16/20 | 80.0 |
| Task 4 | NavigateKitchen | 5 / 5 | 20/20 | 100.0 |
| Task 5 | OpenCabinet | 7 / 7 | 18/20 | 90.0 |
| Task | Name | Layout / Style | Success / Total | Success Rate |
| Task 0 | CloseBlenderLid | 8 / 8 | 6/20 | 30.0 |
| Task 1 | CloseFridge | 2 / 2 | 20/20 | 100.0 |
| Task 2 | CloseToasterOvenDoor | 7 / 7 | 8/20 | 40.0 |
| Task 3 | CoffeeSetupMug | 4 / 4 | 8/20 | 40.0 |
| Task 4 | NavigateKitchen | 5 / 5 | 8/20 | 40.0 |
| Task 5 | OpenCabinet | 7 / 7 | 18/20 | 90.0 |
| DreamZero (WAM) | HarnessPAI ( zero-shot ) | |||
| Task | Swap | Task | Swap | Task |
| task0 (alphabet soup) | 0.0 | 0.0 | 100.0 | 94.0 |
| task1 (cream cheese) | 0.0 | 96.0 | 84.0 | 100.0 |
| task2 (salad dressing) | 0.0 | 0.0 | 100.0 | 100.0 |
| task3 (BBQ sauce) | 0.0 | 0.0 | 84.0 | 96.0 |
| task4 (ketchup) | 0.0 | 0.0 | 96.0 | 78.0 |