cs.ROOct 20, 2025

Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling

Authors: Po-Chen Ko, Yu Chan, Yu-Hsiang Fu, Hsien-Jeng Yeh, Chu-Rong Chen, Dong-Chen Tsai, Wei-Chiu Ma, Yilun Du, +2 more

Organizations: National Taiwan University · MIT CSAIL · Cornell University · Harvard University

Abstract

Video planning has emerged as a flexible framework for robot manipulation, in which a generative model predicts a video of task completion, and a downstream module translates the predicted frames into actions. However, existing methods typically ignore information from past interactions, limiting their ability to adapt to latent physical properties that can only be revealed through trial and error, such as whether a door should be pushed or pulled, or how friction affects object dynamics. When a plan fails, these methods usually replan from scratch without leveraging the information revealed by the failure. To address this limitation, we introduce RELIC, REplanning with Latent embedding refInement and Candidate rejection, a video planning framework that adapts to hidden physical properties from test-time failures. RELIC optimizes a latent embedding that captures the environment's hidden physical properties from interaction videos and introduces a rejection-based sampling mechanism that filters out hypotheses inconsistent with prior failures. Across eight tasks in two simulation suites, RELIC consistently reduces the number of replanning steps required for success, and linear probes show that its embedding captures the hidden parameters from the interaction itself rather than from scene appearance. Across four challenging real-world robotic manipulation tasks involving hidden interaction modes, e.g., friction, center of mass, and object mass, RELIC raises the one-shot replanning success rate of a video planning baseline from 30.0% to 63.8% after a single physical interaction.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback

    Sep 29, 2026Ahmed Nader Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul +2Long-Horizon ManipulationPerception

  2. NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning

    Feb 23, 2026Jiahui Fu, Junyu Nan, Lingfeng Sun +7

  3. Grounded World Model: Latent Planning with Language Goals

    Apr 13, 2026Quanyi Li, Lan Feng, Haonan Zhang +4World ModelsWorld