Organizations: Renmin University of China · Tsinghua University · Shanghai Qizhi Institute · University of Nottingham · Beijing Academy of Artificial Intelligence
Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts inside COACHWORLD, our shared action-conditioned world model, and uses a progress judge to record the first subtask that fails to complete. Aggregated records select which subtask demonstrations to acquire and which expert adapters to update. Across two simulation suites and two real-robot platforms, imagined and deployed success correlate over 22 task-policy pairs (rho = 0.840). Controlled comparisons show that our coaching method outperforms matched baselines under matched data budgets and update schedules. With only 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The coached experts also transfer to four held-out compositions, achieving an average success of 35.0%, compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. Together, these results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://robocoach-ai.github.io/
Figures & tables
Figure 1 : CoachWorld architecture. Sparse visual history, task instructions, and calibrated two-slot end-effector trajectories condition future-video prediction across single-arm and bimanual embodiments.
Method
Evaluation set
LPIPS ↓
FVD ↓
HSD ↑
nDTW ↑
DYN ↑
Phys. Adh. ↑
Instr. Follow. ↑
Ctrl-World
DROID-180
0.2078
102.97
0.2234
0.2323
0.0668
0.6067
0.7867
Lab Franka-180
0.3060
493.20
0.1304
0.1037
0.0458
0.2600
0.2400
Cosmos 3
DROID-180
0.2830
140.62
0.1702
0.1698
0.0571
0.7267
0.5467
Lab Franka-180
0.3397
634.75
0.1658
0.1374
0.0653
0.3900
0.2467
OSCAR-2B
DROID-180
0.2248
101.92
0.1785
0.1838
0.0456
0.5267
0.5467
Lab Franka-180
0.2014
236.78
0.1909
0.2110
0.1412
0.4800
0.5833
Table 1 : Paired full-episode action-conditioned prediction. All models receive matched visual histories and future actions on DROID-180 and Lab Franka-180. The two CoachWorld variants compare mixed- and single-domain training at matched training volume. LPIPS/FVD measure video fidelity, HSD/nDTW/DYN measure end-effector trajectory agreement, and Phys. Adh./Instr. Follow. are mean blind human ratings on a 1–5 scale, divided by five. Best and second-best values within each evaluation set are bold and underlined, respectively.
Figure 2 : Imagined versus deployed policy success. All four platforms share one coordinate system and one pooled OLS fit.
Observation source
Switch MAE (s) ↓
Switch Rec.1.0s (%) ↑
Spearman ρ↑
TP/TN/FP/FN
Out.-Prec. (%) ↑
Out.-Rec. (%) ↑
Out.-F1 (%) ↑
Reference video
0.414
82.73
0.727
101/92/18/9
84.87
91.82
88.21
CoachWorld rollout
0.815
62.73
0.638
96/87/23/14
80.67
87.27
83.84
Table 2 : RoboMeter reliability under reference and world-model observations. We test RoboMeter on reference video and CoachWorld -generated rollout. Switch MAE and Recall@1.0s measure stage-transition timing, Spearman ρ measures terminal-stage rank agreement, and outcome precision, recall, and F1 measure binary transition decisions on adjacent-stage examples.
Figure 3 : Complete-task success over coaching rounds. We evaluate task-averaged success rate as cumulative coaching demonstrations increase. Error bars show 95% confidence intervals for task-averaged success, computed from per-task evaluation counts (Appendix 10 ). As a result, RoboCoach achieves the highest success rate in all four settings.
Condition
Data selection
Update location
Single VLA + Uniform
Uniform
Shared global adapter
Single VLA + WM-targeted
WM-targeted
Shared global adapter
Modular + Random
Random subtask–expert pairs
Selected skill experts
RoboCoach
WM-targeted
Selected skill experts
Table 3 : Acquisition and update conditions.
Figure 4 : Held-out composition case studies. Execution sequences on AgileX (A, C) and Franka (B, D), illustrating longer continuation (A), cross-task composition (B), and skill reordering (C, D). Frames in each sequence progress from left to right.
Figure 5 : Real-robot platforms.
Platform
Task
Ordered subtasks
Franka
Lamp switch and cord
press the lamp switch; pull the lamp cord
Franka
Two buttons press
press the green button; press the red button
Franka
Bread and pan
pick the bread; place the bread into the pan; pick the lid; place the lid over the pan
AgileX
Block and drawer
pick the red block; place the block into the drawer; close the drawer
AgileX
Shoes and box
pick the right shoe; place the right shoe into the shoebox; pick the left shoe; place the left shoe into the shoebox; close the box
AgileX
Table setting
pick the cloth; wipe the table; push the plate to the center of the table; pick the cup; place the cup on the plate
Table 4 : Real-robot task corpus and ordered decompositions.
Evaluation set
Episodes
Min (s)
Q1 (s)
Median (s)
Q3 (s)
Max (s)
DROID-180
180
5.8
10.2
14.2
23.9
95.4
Laboratory Franka-180
180
17.0
33.8
37.6
44.3
66.8
Table 5 : Length distribution of the current full-episode world-model suites. Duration is measured on the 5 Hz evaluation timeline after the initial history.
ID
Task
Demos
Avg. steps
ID
Task
Demos
Avg. steps
L0
Soup + tomato sauce in basket
38
258.1
L5
Book in rear caddy compartment
33
290.0
L1
Cream cheese + butter in basket
36
250.6
L6
White mug on plate + pudding right
29
407.2
L2
Turn on stove + place moka pot
34
293.8
L7
Soup + cream cheese in basket
49
259.2
L3
Black bowl in drawer + close
41
265.0
L8
Both moka pots on stove
35
245.1
L4
Two mugs on left/right plates
43
267.3
L9
Mug in microwave + close
41
186.2
Table 6 : LIBERO-Long task pool. Task IDs L0–L9 are used throughout the per-task coaching results in the appendix. Task names are shortened versions of the original language instructions. The pool contains 379 successful demonstrations.
ID
Task
Avg. steps
Instruction
R0
Blocks Ranking RGB
466
Arrange the red, green, and blue blocks from left to right.
R1
Blocks Ranking Size
466
Arrange the three blocks from largest to smallest.
R2
Put Bottles Dustbin
637
Place the bottles into the dustbin on the left.
R3
Stack Bowls Three
476
Stack the three bowls.
R4
Stack Blocks Three
481
Stack blue on green and green on red.
Table 7 : RoboTwin 2.0 task pool. We use five horizon-3 tasks in the controlled coaching study. Task IDs R0–R4 are used throughout the per-task coaching results in the appendix. Each task contains 50 clean and 500 randomized demonstrations, for 2,750 trajectories in total.
Item
Setting
Initialization
Wan2.2 TI2V-5B
Frequency
5 Hz
Image size
512×768
History:future latents
5:3
Optimizer
AdamW with cosine decay
Learning rate
1×10−5 ; 5×10−6 late stage
Table 8 : CoachWorld training configuration.
DROID-180
Lab Franka-180
Method
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
Ctrl-World
20.597
0.7972
21.859
0.8279
Cosmos 3
15.776
0.5916
18.308
0.6654
OSCAR-2B
17.304
0.7449
21.915
0.8473
CoachWorld (single-domain)
21.115
0.8696
22.193
0.8649
CoachWorld (mixed-domain)
21.438
0.8768
22.168
0.8665
Table 9 : Supplemental full-episode visual metrics. PSNR and SSIM are secondary appearance diagnostics. The CoachWorld variants compare mixed-domain and single-domain training at matched volume.
Figure 6 : Qualitative full-episode prediction. Representative episodes shown at matched points on the recorded action clock.
Figure 7 : Representative CoachWorld rollouts on single-arm domains. Top rows show reference episodes and bottom rows show CoachWorld predictions; six frames uniformly span each full episode.
Figure 8 : Representative CoachWorld rollouts on bimanual domains. Top rows show reference episodes and bottom rows show CoachWorld predictions; six frames uniformly span each full episode.
Judge
Switch MAE (s) ↓
Switch Rec.1.0s (%) ↑
Spearman ρ↑
TP/TN/FP/FN
Out.-Prec. (%) ↑
Out.-Rec. (%) ↑
Out.-F1 (%) ↑
LIV †
3.311
23.64
0.328
64/67/43/46
59.81
58.18
58.99
Contrastive λ †
1.905
28.18
0.603
95/80/30/15
76.00
86.36
80.85
VLM-based
Qwen3-VL
0.893
20.00
0.151
29/82/28/81
50.88
26.36
34.73
TOPReward †
1.603
33.64
0.330
76/55/55/34
58.02
69.09
63.07
SuccessVQA †
0.859
59.09
0.286
82/72/38/28
68.33
74.55
71.30
Table 10 : Progress-judge comparison on causal reference-video replay. Switch metrics evaluate stage-transition timing, Spearman ρ measures terminal-stage rank agreement, and outcome metrics evaluate binary transition decisions on adjacent-stage examples. Best and second-best values are bold and underlined, respectively.
Figure 9 : From round-level diagnosis to target-level coaching evidence. Left: RoboCoach aggregates rollout outcomes into task-level success rates and ranks recurring first-timeout subtasks as coaching targets. Right: selecting a target reveals its supporting episode-level evidence, including the failed rollout, judge progress trace, and router-dispatched subtask sequence. This drill-down connects aggregate failure diagnosis to concrete and auditable coaching examples.
Domain
Method
Round 1
Round 2
Round 3
LIBERO
Random
L6: pick white mug → pick L8: place first moka pot → place
L0: pick tomato sauce → pick L7: place alphabet soup → place
L7: pick cream-cheese box → pick L9: place mug in microwave → place
L0: pick soup / tomato sauce → pick L8: pick first / second moka pot → pick
L8: place first moka pot → place L9: place mug in microwave → place
RoboTwin
Random
R1: pick largest block → pick R3: pick third bowl → pick
R0: pick blue block → pick R3: place first bowl → place
R3: pick second bowl → pick R4: place second block on first → place
RoboTwin
RoboCoach
R1: place medium block → place R2: pick first bottle → pick
R0: pick green block → pick R1: place largest block → place
R4: place third block on second → place R3: place third bowl on second → place
Franka
RoboCoach
pull lamp cord → pull place lid → place
press green button → press place lid → place
pull lamp cord → pull press red button → press
AgileX
RoboCoach
pick right shoe → pick place tea bag → place
pick red block → pick push plate → push
pick red block → pick place left shoe → place
Table 11 : Per-round acquisition trace. RoboCoach rows report the two scorecard-selected subtask–expert targets in each round, while Random rows report the two randomly selected targets. Single VLA + WM-targeted shares the RoboCoach acquisitions; Uniform has no targeted selection.
Domain
Method
SR@0
SR@B
SR@2B
SR@3B
Budget AUC
LIBERO
Single VLA + Uniform
66.0%
66.8%
66.4%
67.0%
66.6%
LIBERO
Single VLA + WM-targeted
66.0%
68.2%
66.8%
67.8%
67.3%
LIBERO
Modular + Random
66.0%
67.8%
69.2%
68.4%
68.1%
LIBERO
RoboCoach
66.0%
69.6%
70.6%
71.2%
69.6%
RoboTwin
Single VLA + Uniform
58.8%
58.4%
58.0%
59.6%
58.5%
RoboTwin
Single VLA + WM-targeted
58.8%
54.8%
57.2%
54.8%
56.3%
Table 12 : Per-round complete-task success. B=100 in simulation and B=50 on hardware.
Task ID
SR@0
SR@B
SR@2B
SR@3B
Single VLA + Uniform
L0
25/50 (50%)
29/50 (58%)
30/50 (60%)
25/50 (50%)
L1
30/50 (60%)
35/50 (70%)
32/50 (64%)
36/50 (72%)
L2
39/50 (78%)
38/50 (76%)
35/50 (70%)
35/50 (70%)
L3
43/50 (86%)
45/50 (90%)
45/50 (90%)
44/50 (88%)
L4
33/50 (66%)
26/50 (52%)
29/50 (58%)
30/50 (60%)
Table 13 : Per-task LIBERO coaching results under the fixed evaluation protocol. Task IDs follow the LIBERO-Long mapping in Table 6 . Each cell reports successes out of 50 fixed-seed evaluation episodes, with the corresponding success rate in parentheses. All methods share the same SR@0 evaluation of 330/500 (66.0%).
Task ID
SR@0
SR@B
SR@2B
SR@3B
Single VLA + Uniform
R0
25/50 (50%)
26/50 (52%)
25/50 (50%)
25/50 (50%)
R1
22/50 (44%)
22/50 (44%)
20/50 (40%)
24/50 (48%)
R2
33/50 (66%)
41/50 (82%)
38/50 (76%)
37/50 (74%)
R3
39/50 (78%)
37/50 (74%)
43/50 (86%)
41/50 (82%)
R4
28/50 (56%)
20/50 (40%)
19/50 (38%)
22/50 (44%)
Table 14 : Per-task RoboTwin 2.0 coaching results under the fixed evaluation protocol. Task IDs follow the RoboTwin task mapping in Table 7 . Each cell reports successes out of 50 fixed-seed evaluation episodes, with the corresponding success rate in parentheses. All methods share the same SR@0 evaluation of 147/250 (58.8%).
Task
SR@0
SR@B
SR@2B
SR@3B
Single VLA + Uniform
Lamp switch & cord
3/20 (15%)
6/20 (30%)
2/20 (10%)
8/20 (40%)
Two buttons
3/20 (15%)
2/20 (10%)
3/20 (15%)
4/20 (20%)
Bread & pan
2/20 (10%)
1/20 (5%)
7/20 (35%)
6/20 (30%)
Total
8/60 (13.3%)
9/60 (15.0%)
12/60 (20.0%)
18/60 (30.0%)
RoboCoach (ours)
Table 15 : Per-task coaching results on Franka. Each entry reports successes out of 20 evaluation trials, with success rate in parentheses.
Task
SR@0
SR@B
SR@2B
SR@3B
Single VLA + Uniform
Block & drawer
12/20 (60%)
14/20 (70%)
14/20 (70%)
12/20 (60%)
Shoes & box
5/20 (25%)
6/20 (30%)
4/20 (20%)
5/20 (25%)
Table setting
13/20 (65%)
16/20 (80%)
17/20 (85%)
17/20 (85%)
Tea making
2/20 (10%)
3/20 (15%)
2/20 (10%)
4/20 (20%)
Total
32/80 (40.0%)
39/80 (48.8%)
37/80 (46.3%)
38/80 (47.5%)
Table 16 : Per-task coaching results on AgileX. Each entry reports successes out of 20 evaluation trials, with success rate in parentheses.
Case
Trials
RoboCoach
Single VLA + Uniform
A: continuation
20
13/20 (65%)
0/20 (0%)
B: composition
20
3/20 (15%)
0/20 (0%)
C: AgileX reordering
20
5/20 (25%)
0/20 (0%)
D: Franka reordering
20
7/20 (35%)
0/20 (0%)
Task macro
–
35.0%
0.0%
Table 17 : Held-out composition success by case. Task-macro success is the arithmetic mean of the four case-level rates.
Post-training is essential for turning pretrained generalist robot policies into reliable task-specific controllers, but existing human-in-the-loop pipelines remain tied to physical execution: each correction requires robot time, scene setup, resets, and operator supervision in the real world. Meanwhile, action-conditioned world models have been studied mainly for imagination, synthetic data generation, and policy evaluation. We propose \textbf{Human-in-the-World-Model (Hi-WM)}, a post-training framework that uses a learned world model as a reusable corrective substrate for failure-targeted policy improvement. A policy is first rolled out in closed loop inside the world model; when the rollout becomes incorrect or failure-prone, a human intervenes directly in the model to provide short corrective actions. Hi-WM caches intermediate states and supports rollback and branching, allowing a single failure state to be reused for multiple corrective continuations and yielding dense supervision around behaviors that the base policy handles poorly. The resulting corrective trajectories are then added back to the training set for post-training. We evaluate Hi-WM on three real-world manipulation tasks spanning both rigid and deformable object interaction, and on two policy backbones. Hi-WM improves real-world success by 37.9 points on average over the base policy and by 19.0 points over a world-model closed-loop baseline, while world-model evaluation correlates strongly with real-world performance (r = 0.953). These results suggest that world models can serve not only as generators or evaluators, but also as effective corrective substrates for scalable robot post-training.
Yaxuan Li, Zhongyi Zhou, Yefei Chen +5
Current Robotics · Tsinghua University · Peking University +1
Video world models are emerging as a scalable alternative for evaluating generalist robot policies, bypassing the physical constraints and engineering burdens of real-world deployment. However, evaluating policies with video world models remains challenging, as world-model errors can make generated rollouts unreliable and slow inference limits large-scale throughput. We introduce RoboWorld, an automated evaluation pipeline that pairs a fast autoregressive video world model with a task-progress-aware vision-language model scoring. To enable reliable long-horizon autoregressive world-model rollouts, we propose Step Forcing, which combines anchored and one-step self-forwarded contexts to reduce train-test mismatch while preserving action-observation dynamics. Together, these components enable RoboWorld to align strongly with real-world robot evaluation across tasks and environments, achieving Pearson's r = 0.989 and Spearman's ρ = 0.970.
Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.
Sicheng Xie, Yitong Chen, Haidong Cao +3
Institute of Trustworthy Embodied AI, Fudan University · Shanghai Innovation Institute · NeoteAI.