Vision-language-action (VLA) policies trained on large datasets are capable within their training domains, yet they still fail to generalize to the variety of situations a robot meets in real-world deployment. Agentic robot systems complement the policy with a vision-language model (VLM) orchestrator that learns when to call the policy, how to instruct it, and when to use scripted skills instead. However, because the harness is built around a frozen policy that has limited language steerability, the orchestrator can avoid the policy's failures but never overcome them. The policy becomes the bottleneck of the whole system. Fine-tuning the policy can remove this bottleneck, but updating it alone decouples it from an orchestrator tuned to its old behavior. We propose Robo-COP, in which the orchestrator and policy co-evolve during deployment. Robo-COP curates skill demonstrations from its own executions, fine-tunes the policy when this data can address recurring failures, and adopts each new policy only after it improves the skills it was trained for. Across ten simulated RoboLab tasks, Robo-COP raises mean held-out success from 64.8% to 73.8% over the same harness with a frozen policy, while fine-tuning on a fixed schedule without verification reaches only 65.8%. On three real-world tasks, Robo-COP raises held-out success from 38.3% to 50.0%. Robo-COP turns deployment into a self-improving flywheel in which robots learn by doing, with each improvement in execution producing better data for the next round of learning. Videos and code are available at https://robo-cop.pages.dev/.
Figures & tables
Figure 1 : Two coupled loops for deployment-time improvement. The inner loop adapts the harness through episode reflection. The outer loop uses collected experience to update the policy, then revises memory so the harness can reassess the policy’s capabilities.
Figure 2 : The Robo-COP deployment-to-learning loop. A: The orchestrator mixes scripted motion and VLA calls, logs one action stream, and updates memory. B: Judges segment the stream into labeled skills; a controller decides when to train. C: Candidates are verified before adoption, and policy-dependent memory is revised.
Figure 3 : Skill-level data curation. Three recorded examples from Recycle cup clutter: both judges accept the pickup (A), the text judge accepts a grasp that the video judge rejects (B), and both reject a placement (C). The retained pickup comes from an episode that ultimately failed.
Figure 4 : Representative RoboLab tasks. Object configurations vary across deployment and held-out evaluation.
Figure 5 : Held-out success on 50 shared initializations per task. Left: per-task success for all five methods. Right: equally weighted ten-task means; whiskers are 95% beta-posterior credible intervals over the fixed task set (Appendix H ).
Regime
Tasks
π0.5
ASPIRE
Frozen
Fixed
Ours
No training
Adopted
Reverted (final)
VLA alone
3
82.0
66.0
89.3
89.3
80.0
1
2
0
Primitives alone
4
2.5
73.5
55.0
50.0
72.0
1
2
1
Neither alone
3
8.0
16.0
53.3
63.3
70.0
1
2
0
Table 1: Held-out success by task regime (%). Tasks are grouped by whether they can be completed using the pretrained VLA alone (Bananas in bin, Bananas out of bin, Biggest fruit), scripted primitives alone (Fill the empty bin, Sort blocks by initial stack position, Return the misplaced object, Restack onto the unstacked can), or neither (Recycle cup clutter, Recycle sort, Fruits to bowl, except orange). The last three columns count tasks by Robo-COP ’s final policy decision.
Figure 6 : Deployment success on Recycle cup clutter. Success rates use a trailing 25-trial window, starting at trial 25, to reduce fluctuations from randomized scenes and stochastic agent decisions and show longer-term deployment trends. Shading marks candidate evaluation; labels mark training requests and adoption decisions.
Figure 7 : Real-world tasks on a DROID setup. Each task requires placing multiple objects, and a trial succeeds only if every condition is met within the time limit shown.
Robo-COP decisions
Learning (50 trials)
Test (20 trials)
Task
Train request
Outcome
Frozen
Robo-COP
Frozen
Robo-COP
Cup and tape ( 40s )
Train@30
Revert
46
44
40
40
Lidded pot ( 60s )
Train@30
Adopt
32
38
35
70
Mustard and carrots ( 90s )
Wait
n/a
32
32
40
40
Mean
36.7
38.0
38.3
50.0
Table 2 : Real-world success rates (%) and Robo-COP ’s training decisions. Frozen uses the same harness but never updates the policy.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Constant
Value
Raising it
Turns per episode
20
more deliberation per episode, fewer episodes per budget
Memory budget
8000 char
more recall, less of the prompt in the attended region
Stillness exit
16 steps
fewer premature exits, more budget spent on stalled bursts
Held frames per segment
15
stronger stopping behaviour, more static frames in D
Action horizon H
15
longer commitment per inference, coarser reaction
Curated floor / new since last
30 / 10
later first training run, more evidence behind it
Appendix
Table 3 : Constants defining the loop.
Task
Trial
Selected target instructions
Data (ep./seg.)
Decision
Bananas in bin
–
–
–
Base retained
Bananas out of bin
45
Pick up the banana; place the banana on the table
43/185
Adopt
Fill the empty bin
–
–
–
Base retained
Sort blocks by initial stack position
60
Place the green block in the red bowl; place the blue block on the table; pick up the red block
59/280
Adopt
Return the misplaced object
60
Pick up the blue can; place the blue can in the grey bin
58/220
Revert → Adopt
Recycle cup clutter
45
Pick up the coffee cup; place the coffee cup in the black bin
43/194
Adopt → Adopt
Appendix
Table 4 : First training request and later decisions by task. Trial, targets, and data (curated episodes / kept segments) refer to the first training request; arrows show later decisions. Six tasks end with an updated policy. A dash means the controller never requested training.
Arm
Closest setting
Actuation
Adaptation
Deployment
π0.5 only
base VLA
policy
none
eval only
Frozen-policy baseline
memory-adaptive harness
both
memory
100 trials
Fixed schedule
scheduled policy update
both
memory+policy
100 trials
ASPIRE
agentic skill discovery
programs
skill library
100 search budget
Robo-COP
full method
both
memory+policy
100 trials
Appendix
Table 5 : Simulation comparison arms. Frozen, Fixed, and Robo-COP have matched 100-trial deployment sequences; all methods are tested on held-out initializations.
Deployment
Held-out
Task
Frozen
Fixed
Ours
π0.5
Frozen
Fixed
ASPIRE
Ours
Bananas in bin
80
88
91
100
88
98
88
92
Bananas out of bin
69
74
66
74
90
88
68
72
Fill the empty bin
96
84
94
8
90
82
98
80
Sort blocks by initial stack position
61
62
58
0
6
42
72
60
Return the misplaced object
69
71
82
0
92
50
70
86
Appendix
Table 6 : Simulation success rates (%). Deployment uses 100 trials per task and held-out evaluation uses 50. ASPIRE is evaluated on held-out trials only. Bold marks the best result in each group.
Final k trials
10
20
25
30
50
Frozen-policy baseline
69.0
63.5
67.6
68.7
66.2
Fixed schedule
68.0
65.5
67.2
68.7
68.4
Robo-COP
75.0
75.5
73.6
72.7
73.4
Appendix
Table 7 : Deployment success (%) over the final k trials.
Figure 8 : Trailing 10-episode deployment success on each task.
Figure 9 : Held-out task differences. Bars show Robo-COP minus each baseline, in percentage points. The zero for Fruits to bowl, except orange versus fixed schedule is a tie. The lower panel shows the range of the mean gap when one task is omitted; dots mark the full ten-task means.
Figure 10 : Successful episodes for tasks 1–5. Each row shows five frames from one trial, at roughly 2, 25, 50, 75, and 98% of its duration.
Figure 11 : Successful episodes for tasks 6–10. Frames are sampled as in Figure 10 .
Figure 12 : Initial object positions after 50 seeded resets per task, in simulator world coordinates (m). Colors distinguish objects within each task.
Figure 13 : Successful real-world executions of the three tasks. Frames are shown in temporal order.
While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail. In this work, we propose ARCHITECT, a framework that treats robot policy acquisition as an interactive program synthesis task. ARCHITECT leverages the reasoning capabilities of LLM coding agents to synthesize modular robot programs that utilize a suite of perception and control tools. Unlike end-to-end models where distribution shift leads to unpredictable, cascading failures, our modular architecture allows users to isolate failures and localize feedback at the level of abstraction required. We introduce an iterative process where a human supervisor provides natural language corrections to steer the policy. These corrections are grounded in the policy code by program execution traces and distilled into a persistent skill library, a form of long-term in-context learning which enables the agent to accumulate a repertoire of reusable, interpretable behaviors. In a benchmark evaluation on a Franka Panda robot, ARCHITECT outperforms state-of-the-art VLA models and program synthesis baselines on complex, long-horizon tasks, including articulated object manipulation and cloth folding. Our results demonstrate that the synthesized skill library enables the system to transfer to novel tasks with decreasing human intervention, providing a steerable and data-efficient alternative to black-box robot learning. Website: https://robo-architect.github.io/
Daphne Chen, Archit Ritesh Jain, Eric Goossen +4
University of Washington · 2Microsoft Research · 3Massachusetts Institute of Technology
While Vision-Language-Action (VLA) models demonstrate impressive capabilities in robotic manipulation, their memoryless nature renders them brittle to test-time environment shifts, particularly hardware shifts caused by wear or imperfect calibration. Enabling these models to self-adapt during deployment without requiring continuous on-site recalibration remains a critical bottleneck for real-world scalability. In this work, we introduce Self-Adaptive VLA, a novel post-training recipe that enables the policy to iteratively adapt to deployment-time hardware shifts leveraging its own rollouts as context. To do so, we first collect policy rollouts under deliberately injected hardware shifts. We then transform the base policy's training data into shift-conditioned expert demonstrations by pre-compensating the expert actions for these known shifts. Next, we introduce a lightweight, plug-in context encoder that compresses the context, including visual observation, proprioception, and actions in the shifted environment, into a latent context token. This token modulates the policy through adaptive layer normalization (AdaLN). Furthermore, we find that context tokens can be ensembled, allowing the policy to iteratively self-correct and mitigate failures step by step. Extensive experiments across four precision-critical bi-manual and dexterous manipulation tasks show that Self-Adaptive VLA recovers over 80% of the base policy's performance under hardware shifts, such as actuation bias and joint encoder offsets. Moreover, Self-Adaptive VLA enables more robust deployment to new workstations compared to the base policy. Our approach provides a pathway for robust large-scale real-world robot deployments and easier maintenance. See videos at https://icefoxzhx.github.io/self-adaptive-vla.
Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose FailPatch, a failure-driven residual patching framework that decouples action supervision from execution-reliability supervision. Successful demonstrations ground how the policy should act, while deployment trajectories indicate when its behavior becomes unreliable. We further observe that action hidden representations exhibit clear linear separability between reliable and failure-associated states while directly conditioning action generation. Building on these insights, FailPatch introduces a Null-gated Residual Expert Bank into the action hidden space of a frozen VLA policy. A unified Preserve--Redirect--Trust objective retains the original policy in reliable states, selects residual experts in failure-associated states and redirects representations from failure regions toward success-associated regions under bounded intervention. With only 0.52% trainable parameters, FailPatch improves success rates by 11.0 percentage points on four long-horizon RoboTwin tasks under clean evaluation, 9.5 percentage points under clean-to-random generalization, and 16.7 percentage points over the baseline across three real-world tasks. Project and code: https://github.com/yupeng-2003/FailPatch.
Peng Yu, Jiacheng Wang, Ziheng Zhang +6
Xi’an Jiaotong University · Dexmal · Nanjing University