Video planning has emerged as a flexible framework for robot manipulation, in which a generative model predicts a video of task completion, and a downstream module translates the predicted frames into actions. However, existing methods typically ignore information from past interactions, limiting their ability to adapt to latent physical properties that can only be revealed through trial and error, such as whether a door should be pushed or pulled, or how friction affects object dynamics. When a plan fails, these methods usually replan from scratch without leveraging the information revealed by the failure. To address this limitation, we introduce RELIC, REplanning with Latent embedding refInement and Candidate rejection, a video planning framework that adapts to hidden physical properties from test-time failures. RELIC optimizes a latent embedding that captures the environment's hidden physical properties from interaction videos and introduces a rejection-based sampling mechanism that filters out hypotheses inconsistent with prior failures. Across eight tasks in two simulation suites, RELIC consistently reduces the number of replanning steps required for success, and linear probes show that its embedding captures the hidden parameters from the interaction itself rather than from scene appearance. Across four challenging real-world robotic manipulation tasks involving hidden interaction modes, e.g., friction, center of mass, and object mass, RELIC raises the one-shot replanning success rate of a video planning baseline from 30.0% to 63.8% after a single physical interaction.
Figures & tables
Fig. 1 : Video replanning from test-time failures. Top: existing video planning methods condition only on the first frame ( red ). RELIC also stores failed interactions and plans ( blue ) and uses them to generate the next plan. Bottom: a real-robot example. The bar’s center of mass is hidden, so the first plan fails; the failure reveals it, and RELIC replans off-center.
Meta-World
ManiSkill3
Method
PushBar
PickBar
SlideBrick
OpenBox
TurnFaucet
Seesaw
Bounce
Strike
Overall [-2pt] (all 8, norm.)
CQL
9.96± 0.20
9.22± 0.19
8.94± 0.18
3.32± 0.18
5.63± 0.21
5.91± 0.18
6.03± 0.68
4.96± 0.37
2.12± 0.08
BC
10.26± 0.30
8.79± 0.32
9.24± 0.33
1.59± 0.19
3.54± 0.27
5.17± 0.09
7.48± 0.71
4.02± 0.20
1.74± 0.06
AVDC
6.83± 0.27
4.41± 0.22
7.36± 0.27
1.82± 0.16
1.67± 0.16
5.62± 0.47
5.80± 0.48
4.13± 0.76
1.29± 0.06
VLM verifier
4.73± 0.40
4.26± 0.38
7.35± 0.45
1.56± 0.36
1.41± 0.16
5.72± 0.48
5.64± 0.43
4.12± 0.41
1.17± 0.06
RELIC (Refine FS)
4.66± 0.23
3.87± 0.21
6.72± 0.26
1.73± 0.16
1.54± 0.15
5.31± 0.37
5.32± 0.38
2.96± 0.26
1.10± 0.04
TABLE I : Replanning performance across our two simulation task suites ( ↓ ). Per-task columns report raw replanning counts as mean ± SEM over 400 trials; CQL also averages three policies. Overall reports the mean ratio of each method’s replanning count to that of RELIC across the eight tasks, with equal weight for each task; its SEM is propagated from the per-task SEMs by the delta method, assuming independence across methods.
Suite
Tasks
Videos
Successful
Meta-World
5
14,270
670
ManiSkill3
3
18,750
750
TABLE II : Offline experience dataset D per suite. Total number of interaction videos and the number of successful ones; a video is successful when the guessed value equals the true value. Each video is one scripted execution in which the policy assumes a guessed parameter value while the environment uses a true hidden value.
Method
PSNR ↑
LPIPS ↓
CLIP ↑
DINO ↑
AVDC
19.426
0.271
0.870
0.705
RELIC (R3M)
19.632
0.261
0.875
0.705
RELIC (CLIP)
19.584
0.262
0.871
0.707
RELIC (DINOv2, Ours)
19.817
0.255
0.882
0.717
TABLE III : Evaluation on replanned videos. Similarity of replanned videos to ground-truth videos from the scripted policy. Arrows indicate better performance; best values are bold.
Fig. 6 : Ablation studies. We evaluate variations of our method to assess the effectiveness of each component. The number of replans is normalized and aggregated across five tasks.
Latent-conditioned policy
Latent in RELIC pipeline
Task
HiP-MDP
PEARL
HiP-MDP
PEARL
RELIC
PushBar
8.65± 0.24
8.64± 0.23
7.88± 0.24
5.73± 0.21
4.26± 0.20
PickBar
7.65± 0.23
8.99± 0.24
9.62± 0.23
5.70± 0.21
3.84± 0.18
SlideBrick
11.29± 0.22
13.20± 0.13
10.35± 0.23
7.69± 0.23
6.69± 0.26
OpenBox
13.01± 0.15
13.09± 0.14
13.21± 0.13
1.60± 0.04
1.25± 0.10
TurnFaucet
13.03± 0.14
2.57± 0.09
13.06± 0.14
1.60± 0.05
1.39± 0.14
TABLE IV : Comparison with latent-context adaptation methods on the Meta-World suite ( ↓ ). Mean ± SEM replanning counts over 400 trials. The first pair of columns evaluates each method as a latent-context-conditioned actor-critic policy trained with dense rewards and action labels, which RELIC does not use. The second pair replaces our retrieval with the latent inferred by each method and keeps the rest of the pipeline unchanged. Overall is the mean per-task ratio to RELIC with its SEM, as in Table I .
Task
Hidden θ
Metric
Post
Pre
PushBar
center of mass
R2
0.77
−0.002
PickBar
center of mass
R2
0.82
−0.002
SlideBrick
friction
R2
0.84
−5.5
OpenBox
push vs. pull
acc. (%)
100
65
TurnFaucet
CW vs. CCW
acc. (%)
92.5
55
TABLE V : Linear probes decoding the hidden parameter θ from state embeddings. Five-fold cross-validation grouped by configuration ID, out-of-fold metrics. Continuous θ reports R2 ; discrete θ reports accuracy. Post probes embeddings of the interaction video; Pre probes the same encoder’s features of the first frame, before any physical contact.
Fig. 7 : Success rate on PushBar as the entries of D nearest to the test θ are excluded. NN-replay copies the action of the most similar episode retrieved from D . 400 trials per point, 16-attempt cap.
Method
OpenDoor
SlideBlock
Balance
Pinball
Overall
AVDC
8/20
4/20
7/20
5/20
30.0%
RELIC (Ours)
14/20
7/20
17/20
13/20
63.8%
TABLE VI : Real-world success rates. Successes out of trials for each task; Overall reports the macro-averaged success rate. RELIC adapts to hidden interaction modes from a small number of trials.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 9 : Settings of five task environments.
Fig. 15 : Real-world task setups. For each task, we show the overall experimental setup and a detailed view of the task-specific physical design. Dimensions shown in the detail views are reported in millimeters. The Open Door detail is a schematic illustration of the task mechanism and is not drawn to scale. The hidden parameters correspond to push-versus-pull direction for OpenDoor , block friction for SlideBlock , block center of mass for Balance , and ball mass for Pinball .
Task
Hidden parameter
Robot
Trials (S/F)
Open Door
Push vs. pull direction
WidowX (Mobile ALOHA)
10/10
Slide Block
Block friction coefficient
Piper-X
5/20
Balance
Bar’s center of mass
Piper-X
6/12
Pinball
Ball mass
Piper-X
6/12
Appendix
TABLE VII : Real-world task setup. Each task contains a hidden physical parameter that the robot must infer through interaction. Data is collected via tele-operation, reported as successes/failures.
σ
Success Rate
0.0
4/10
0.1 (default)
7/10
0.3
3/10
Appendix
TABLE VIII : Noise injection ablation on Pinball. Success rates with different pixel-level Gaussian noise scales σ applied to the first-frame observation at inference time.
num_parameters
125M
diffusion_resolution
(32, 32)
target_resolution
(128, 128)
base_channels
128
num_res_block
2
attention_resolutions
(2, 4, 8)
channel_mult
(1, 2, 3, 4)
Appendix
TABLE IX : Hyperparameters. Comparison of configuration parameters for the Meta-World benchmark.
PushBar
PickBar
SlideBrick
Method
SR (%) ↑
Trials ↓
Fails ↓
SR (%) ↑
Trials ↓
Fails ↓
SR (%) ↑
Trials ↓
Fails ↓
Latest failure only(Ours)
95.8
4.28 ± 0.20
17
98.0
3.46 ± 0.17
8
82.8
6.64 ± 0.54
69
Full failure history
98.0
4.08 ± 0.19
8
99.0
3.46 ± 0.17
4
77.2
7.39 ± 0.51
91
Appendix
TABLE X : Latest failure versus full interaction history. We report the success rate (SR), the average number of replanning trials (Trials, mean ± SEM), and fail count (Fails) over 400 trials per setting, with refine_steps =100 and M =14 . The best result for each metric within each task is shown in bold.
Fig. 16 : Hidden-state optimization performance over replanning trials. Normalized prediction error over replanning trials. Error is measured on the center-of-mass location for PickBar and PushBar , and on the target push height for SlideBrick . Shaded regions denote 95% confidence intervals over 400 trials. Lower is better.
TABLE XI : Computation time (in seconds) averaged over 100 trials.
PushBar
PickBar
SlideBrick
OpenBox
TurnFaucet
Total
Main exp (100%)
6000/250
6000/250
1560/130
20/20
20/20
13620/650
28%
1000/250
1000/250
1300/130
20/20
20/20
3360/650
Appendix
TABLE XII : Number of demonstrations used for each task. Each cell shows failed/successful demonstrations.
TABLE XIII : Scaling effect on performance. Reported metric is Replans until Success ( ↓ , lower is better). We compare performance with 28% vs. 100% data across each environment.
Fig. 18 : Overall scaling effect. Reported metric is Replans until Success ( ↓ lower is better).
Method \ Dataset Size
(Successful, Failed)
(2, 14)
(10, 70)
(40, 280)
(120, 840)
Random
11.87 ± 1.28
11.87 ± 1.28
11.87 ± 1.28
11.87 ± 1.28
Ours (Retrieval)
18.59 ± 1.54
30.94 ± 1.83
53.12 ± 1.97
67.66 ± 1.85
Appendix
TABLE XIV : Data scaling effect on the 3-faucet task. The table shows the exact match success rate between the hidden parameter of the current environment and that of the retrieved object embedding.
Fig. 19 : Multi-faucet system identification setup. A robot interacts with one of multiple faucet instances that share visual structure but differ in hidden physical properties (e.g., joint stiffness, friction, or rotation direction). The agent must infer the correct interaction mode through probing actions rather than relying on appearance alone.
Method
PushBar
PickBar
SlideBrick
OpenBox
TurnFaucet
Overall (Normalized)
Attract
5.55±0.55
4.16±0.45
12.26±0.52
1.31±0.26
1.31±0.29
1.07
Reject (Ours)
5.95±0.55
4.35±0.47
10.60±0.58
1.13±0.25
1.14±0.27
1.00
Appendix
TABLE XV : Attract vs. Reject strategies.
Fig. 20 : Robustness to task-irrelevant visual noise. We introduce appearance-randomized environments with visually distracting backgrounds while preserving identical interaction dynamics. This setting isolates whether distance functions respond to task-relevant motion rather than background variation.
Method
SlideBrick Noisy
OpenBox Noisy
No Rejection
6.09
1.91
DINO
5.56
1.65
L2 error
4.89
1.79
Seg + DINO
5.54
1.65
Seg + L2 error
4.53
1.65
Appendix
TABLE XVI : Performance under visually noisy environments. Metric: Replans until success.
Fig. 21 : Appearance variation under fixed interaction dynamics. Object color and texture are randomized while geometry and physical properties remain unchanged. This setting decouples visual appearance from interaction behavior and tests adaptation under perceptual aliasing.
Method
PushBar
OpenBox
PickBar
TurnFaucet
SlideBrick
Random Guess
.0954±.0052
.4900±.0500
.0954±.0052
.4900±.0500
.0528±.0040
RELIC-Seen
.0694±.0057
.0200±.0140
.0793±.0069
.0200±.0140
.0448±.0032
RELIC-Unseen
.0693±.0049
.0700±.0255
.0729±.0054
.0200±.0140
.0428±.0032
Appendix
TABLE XVII : Raw retrieval error on hidden system parameters (lower is better).
Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle in open-world conditions, as they are hand-tuned for specific scenarios and lack generalization. Vision-Language Models (VLMs) offer a promising alternative as they combine broad world knowledge with unified visual--text reasoning, enabling them to generalize across diverse scenarios and generate accurate, grounded task plans. However, for effective deployment in dynamic real-world settings, VLMs must be embedded into frameworks capable of handling uncertainty and environmental changes. Existing frameworks broadly address this reactively, triggering replanning only after execution failures or post-task checks, risking failed actions. Some methods verify conditions before actions, but these discrete checks miss changes occurring during execution. To address this, we present ProAct-VLM, an adaptive, physically grounded task planning framework that integrates VLMs within a real-time perception--feedback loop. ProAct-VLM continuously monitors the environment and re-plans as soon as relevant changes are detected, enabling adaptation before failure occurs. Evaluations against multiple baselines and across different VLM backbones show that our framework improves both success rates and efficiency in dynamic, long-horizon manipulation tasks. Project page: https://github.com/moured/ProAct-VLM
Ahmed Nader Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul +2
Khalifa University, Abu Dhabi, United Arab Emirates. · Sereact GmbH, Stuttgart, Germany.
Solving complex long-horizon robotic tasks requires joint reasoning over abstract task structure and low-level physical interaction. While combining Vision-Language Models (VLMs) and video generation models offers a promising path for zero-shot planning, their individual tendencies to hallucinate physics or violate geometric consistency often compound over time, preventing reliable real-world execution. We introduce NovaPlan, a hierarchical framework that enables robust, zero-shot long-horizon manipulation by systematically proposing, verifying, and repairing visual plans. At the high level, a VLM planner decomposes tasks and filters out dynamically inconsistent futures by verifying multiple candidate video rollouts. To translate these imagined futures into reliable physical actions, NovaPlan utilizes a hybrid geometric representation that adaptively switches between object-centric flow and human hand flow. Finally, NovaPlan closes the loop by continuously monitoring execution to verify outcomes and synthesize local, non-prehensile corrective behaviors, such as fingertip poking, when failures occur. Across diverse multi-stage tasks, NovaPlan substantially outperforms prior zero-shot systems, achieving complex assembly and dexterous error recovery entirely without task-specific training or demonstrations. Please visit our project website for additional results: https://nova-plan.github.io/
Jiahui Fu, Junyu Nan, Lingfeng Sun +7
Robotics and AI Institute · Carnegie Mellon University · Brown University +1
World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model. The frozen readout of this embedding model maps this imagined future and the task description into the same embedding space, where their negative cosine similarity serves as the planning cost. Training GWM requires only offline and task-agnostic video-action pairs and no language labels. In simulated experiments on WISER, planning with GWM, which executes the candidate action of lowest cost, solves 87% of 288 tasks with unseen instructions and visual signals, while ten fine-tuned VLAs average 22%. We then scale GWM up with real robot data, and use it for zero-shot planning in realistic simulation and real scenes. In the IsaacSim evaluation, planning with GWM completes all 70 trials across 14 tasks that require reasoning over referring expressions, matching a modular planner grounded by a frontier VLM, while pi0.5 reaches 37/70. Deployed on a real Franka, the same stack completes 55/60 separately evaluated pick-and-place sub-tasks, comparable to the modular planner's 52/60, with the full system running locally on a single consumer GPU. Project website: https://quanyili.github.io/gwm-wiser/.