Vision-language-action (VLA) policies trained on large datasets are capable within their training domains, yet they still fail to generalize to the variety of situations a robot meets in real-world deployment. Agentic robot systems complement the policy with a vision-language model (VLM) orchestrator that learns when to call the policy, how to instruct it, and when to use scripted skills instead. However, because the harness is built around a frozen policy that has limited language steerability, the orchestrator can avoid the policy's failures but never overcome them. The policy becomes the bottleneck of the whole system. Fine-tuning the policy can remove this bottleneck, but updating it alone decouples it from an orchestrator tuned to its old behavior. We propose Robo-COP, in which the orchestrator and policy co-evolve during deployment. Robo-COP curates skill demonstrations from its own executions, fine-tunes the policy when this data can address recurring failures, and adopts each new policy only after it improves the skills it was trained for. Across ten simulated RoboLab tasks, Robo-COP raises mean held-out success from 64.8% to 73.8% over the same harness with a frozen policy, while fine-tuning on a fixed schedule without verification reaches only 65.8%. On three real-world tasks, Robo-COP raises held-out success from 38.3% to 50.0%. Robo-COP turns deployment into a self-improving flywheel in which robots learn by doing, with each improvement in execution producing better data for the next round of learning. Videos and code are available at https://robo-cop.pages.dev/.
Figures & tables
Figure 1 : Two coupled loops for deployment-time improvement. The inner loop adapts the harness through episode reflection. The outer loop uses collected experience to update the policy, then revises memory so the harness can reassess the policy’s capabilities.
Figure 2 : The Robo-COP deployment-to-learning loop. A: The orchestrator mixes scripted motion and VLA calls, logs one action stream, and updates memory. B: Judges segment the stream into labeled skills; a controller decides when to train. C: Candidates are verified before adoption, and policy-dependent memory is revised.
Figure 3 : Skill-level data curation. Three recorded examples from Recycle cup clutter: both judges accept the pickup (A), the text judge accepts a grasp that the video judge rejects (B), and both reject a placement (C). The retained pickup comes from an episode that ultimately failed.
Figure 4 : Representative RoboLab tasks. Object configurations vary across deployment and held-out evaluation.
Figure 5 : Held-out success on 50 shared initializations per task. Left: per-task success for all five methods. Right: equally weighted ten-task means; whiskers are 95% beta-posterior credible intervals over the fixed task set (Appendix H ).
Regime
Tasks
π0.5
ASPIRE
Frozen
Fixed
Ours
No training
Adopted
Reverted (final)
VLA alone
3
82.0
66.0
89.3
89.3
80.0
1
2
0
Primitives alone
4
2.5
73.5
55.0
50.0
72.0
1
2
1
Neither alone
3
8.0
16.0
53.3
63.3
70.0
1
2
0
Table 1: Held-out success by task regime (%). Tasks are grouped by whether they can be completed using the pretrained VLA alone (Bananas in bin, Bananas out of bin, Biggest fruit), scripted primitives alone (Fill the empty bin, Sort blocks by initial stack position, Return the misplaced object, Restack onto the unstacked can), or neither (Recycle cup clutter, Recycle sort, Fruits to bowl, except orange). The last three columns count tasks by Robo-COP ’s final policy decision.
Figure 6 : Deployment success on Recycle cup clutter. Success rates use a trailing 25-trial window, starting at trial 25, to reduce fluctuations from randomized scenes and stochastic agent decisions and show longer-term deployment trends. Shading marks candidate evaluation; labels mark training requests and adoption decisions.
Figure 7 : Real-world tasks on a DROID setup. Each task requires placing multiple objects, and a trial succeeds only if every condition is met within the time limit shown.
Robo-COP decisions
Learning (50 trials)
Test (20 trials)
Task
Train request
Outcome
Frozen
Robo-COP
Frozen
Robo-COP
Cup and tape ( 40s )
Train@30
Revert
46
44
40
40
Lidded pot ( 60s )
Train@30
Adopt
32
38
35
70
Mustard and carrots ( 90s )
Wait
n/a
32
32
40
40
Mean
36.7
38.0
38.3
50.0
Table 2 : Real-world success rates (%) and Robo-COP ’s training decisions. Frozen uses the same harness but never updates the policy.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Constant
Value
Raising it
Turns per episode
20
more deliberation per episode, fewer episodes per budget
Memory budget
8000 char
more recall, less of the prompt in the attended region
Stillness exit
16 steps
fewer premature exits, more budget spent on stalled bursts
Held frames per segment
15
stronger stopping behaviour, more static frames in D
Action horizon H
15
longer commitment per inference, coarser reaction
Curated floor / new since last
30 / 10
later first training run, more evidence behind it
Appendix
Table 3 : Constants defining the loop.
Task
Trial
Selected target instructions
Data (ep./seg.)
Decision
Bananas in bin
–
–
–
Base retained
Bananas out of bin
45
Pick up the banana; place the banana on the table
43/185
Adopt
Fill the empty bin
–
–
–
Base retained
Sort blocks by initial stack position
60
Place the green block in the red bowl; place the blue block on the table; pick up the red block
59/280
Adopt
Return the misplaced object
60
Pick up the blue can; place the blue can in the grey bin
58/220
Revert → Adopt
Recycle cup clutter
45
Pick up the coffee cup; place the coffee cup in the black bin
43/194
Adopt → Adopt
Appendix
Table 4 : First training request and later decisions by task. Trial, targets, and data (curated episodes / kept segments) refer to the first training request; arrows show later decisions. Six tasks end with an updated policy. A dash means the controller never requested training.
Arm
Closest setting
Actuation
Adaptation
Deployment
π0.5 only
base VLA
policy
none
eval only
Frozen-policy baseline
memory-adaptive harness
both
memory
100 trials
Fixed schedule
scheduled policy update
both
memory+policy
100 trials
ASPIRE
agentic skill discovery
programs
skill library
100 search budget
Robo-COP
full method
both
memory+policy
100 trials
Appendix
Table 5 : Simulation comparison arms. Frozen, Fixed, and Robo-COP have matched 100-trial deployment sequences; all methods are tested on held-out initializations.
Deployment
Held-out
Task
Frozen
Fixed
Ours
π0.5
Frozen
Fixed
ASPIRE
Ours
Bananas in bin
80
88
91
100
88
98
88
92
Bananas out of bin
69
74
66
74
90
88
68
72
Fill the empty bin
96
84
94
8
90
82
98
80
Sort blocks by initial stack position
61
62
58
0
6
42
72
60
Return the misplaced object
69
71
82
0
92
50
70
86
Appendix
Table 6 : Simulation success rates (%). Deployment uses 100 trials per task and held-out evaluation uses 50. ASPIRE is evaluated on held-out trials only. Bold marks the best result in each group.
Final k trials
10
20
25
30
50
Frozen-policy baseline
69.0
63.5
67.6
68.7
66.2
Fixed schedule
68.0
65.5
67.2
68.7
68.4
Robo-COP
75.0
75.5
73.6
72.7
73.4
Appendix
Table 7 : Deployment success (%) over the final k trials.
Figure 8 : Trailing 10-episode deployment success on each task.
Figure 9 : Held-out task differences. Bars show Robo-COP minus each baseline, in percentage points. The zero for Fruits to bowl, except orange versus fixed schedule is a tie. The lower panel shows the range of the mean gap when one task is omitted; dots mark the full ten-task means.
Figure 10 : Successful episodes for tasks 1–5. Each row shows five frames from one trial, at roughly 2, 25, 50, 75, and 98% of its duration.
Figure 11 : Successful episodes for tasks 6–10. Frames are sampled as in Figure 10 .
Figure 12 : Initial object positions after 50 seeded resets per task, in simulator world coordinates (m). Colors distinguish objects within each task.
Figure 13 : Successful real-world executions of the three tasks. Frames are shown in temporal order.