Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle in open-world conditions, as they are hand-tuned for specific scenarios and lack generalization. Vision-Language Models (VLMs) offer a promising alternative as they combine broad world knowledge with unified visual--text reasoning, enabling them to generalize across diverse scenarios and generate accurate, grounded task plans. However, for effective deployment in dynamic real-world settings, VLMs must be embedded into frameworks capable of handling uncertainty and environmental changes. Existing frameworks broadly address this reactively, triggering replanning only after execution failures or post-task checks, risking failed actions. Some methods verify conditions before actions, but these discrete checks miss changes occurring during execution. To address this, we present ProAct-VLM, an adaptive, physically grounded task planning framework that integrates VLMs within a real-time perception--feedback loop. ProAct-VLM continuously monitors the environment and re-plans as soon as relevant changes are detected, enabling adaptation before failure occurs. Evaluations against multiple baselines and across different VLM backbones show that our framework improves both success rates and efficiency in dynamic, long-horizon manipulation tasks. Project page: https://github.com/moured/ProAct-VLM
Figures & tables
Fig. 1: Overview of the proposed ProAct-VLM framework. (a) The system begins with a high-level user prompt, which is then converted into a planning prompt. (b) The VLM generates an initial plan based on this input. (e) Meanwhile,the Environment State Estimator continuously constructs a structured scene representation (object IDs, descriptions, poses, and the current Augmented image). (c) The Plan Monitor tracks this state and identifies task-relevant changes (e.g., object addition, removal, or goal modification). When such changes occur, (f) a re-planning prompt is issued to the VLM. (d) The Motion Planner executes the current plan until re-planning is triggered, enabling adaptive task execution in dynamic environments..
Fig. 2: Overview of the proposed perception pipeline. The system operates in three cyclic phases: Initialization (a), Tracking (b), and Environment Updating (c). In (a) Grounded-SAM detects objects from an input image using text prompts, and generates labeled segmentation masks. These masks are linked to CAD models and processed by FoundationPose for initial 6-DoF pose estimation. During (b), DeAOT tracks object masks across frames, while FoundationPose perform pose tracking for already registered objects. In (c), object detection is re-invoked periodically to identify new objects using a Comparing Mask Results (CMR) strategy, followed by pose estimation and reference update.
Fig. 3: illustrates of E in action. The framework detects, segments, and estimates object poses as the scene evolves, correctly handling object additions and removals.
Fig. 4: Disturbance types used in experiments: (a) object addition, (b) object removal, (c) goal change.
Fig. 5: Prompt templates used by VLM-Replan . The components in orange are specific to the LLM3+Image version, while the blue are specific to our Augmented version, the red text represents what we add to make the prompt for re-planning.
TABLE I: Defined disturbance scenarios used in our evaluation.
Scenario
GPT-4o
Gemini 2.5
Claude 4
Qwen 2.5
Llama 4
Succ
Eff
Succ
Eff
Succ
Eff
Succ
Eff
Succ
Eff
S1
100
100
100
100
100
100
20
100
100
100
S2
100
100
100
100
40
100
70
100
60
97
S3
80
80
100
100
0
0
60
70
100
100
S4
80
92.5
100
100
0
0
40
80
10
50
S5
100
96.6
100
100
0
0
50
90
100
96
TABLE II: Results of Ablation Study
Model
Prompt
S1
S2
S3
S4
S5
S6
S7
Avg
GPT-4o
LLM3+Image
90
40
80
30
100
60
30
61.4
OWG
100
80
70
70
100
90
70
83
Ours
100
100
80
80
100
90
70
89
VILA/REPLAN-VLM
100
0
100
40
80
30
0
50
Gemini 2.5 Pro
LLM3+Image
100
90
100
100
100
100
100
98.6
OWG
100
100
90
100
100
100
100
98.6
TABLE III: Results of Comparative Studies .
Fig. 6: The robot executes an initial VLM-generated plan; after a new object is introduced and detected by E (screenshot 8), the plan monitor triggers re-planning. The updated plan is executed successfully (screenshots 9–12).
Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints such as limited free space and object collisions. In this work, we propose APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning. APIVOT learns to leverage language for semantic reasoning, while using visual thoughts as imagined future states for internal verification of geometric feasibility. On long-horizon kitchen tasks, APIVOT outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings. We find that APIVOT learns meaningful modality selection behavior, demonstrating that adaptive interleaving of vision-language thoughts improves both planning success and reasoning efficiency.
Long-horizon manipulation remains challenging for vision-language-action (VLA) policies: real tasks are multi-step, progress-dependent, and brittle to compounding execution errors. We present LoHo-Manip, a modular framework that scales short-horizon VLA execution to long-horizon instruction following via a dedicated task-management VLM. The manager is decoupled from the executor and is invoked in a receding-horizon manner: given the current observation, it predicts a progress-aware remaining plan that combines (i) a subtask sequence with an explicit done + remaining split as lightweight language memory, and (ii) a visual trace -- a compact 2D keypoint trajectory prompt specifying where to go and what to approach next. The executor VLA is adapted to condition on the rendered trace, thereby turning long-horizon decision-making into repeated local control by following the trace. Crucially, predicting the remaining plan at each step yields an implicit closed loop: failed steps persist in subsequent outputs, and traces update accordingly, enabling automatic continuation and replanning without hand-crafted recovery logic or brittle visual-history buffers. Extensive experiments spanning embodied planning, long-horizon reasoning, trajectory prediction, and end-to-end manipulation in simulation and on a real Franka robot demonstrate strong gains in long-horizon success, robustness, and out-of-distribution generalization. Project page: https://www.liuisabella.com/LoHoManip
Building generalizable agents for diverse applications remains a fundamental challenge. While imitation learning-based policies succeed in specific training environments, they often fail to generalize to novel scenes and tasks. In this work, we propose World Action Planner, a robot planning system that leverages the reasoning capabilities of Vision-Language Models (VLMs) and the physical grounding of a multi-task pose-image conditioned world model. Our system enables an agent to propose initial action plans and iteratively refine them via optimization and search, reasoning over imagined world model rollouts. We demonstrate that our approach achieves superior performance across compositional tasks, new layouts, and zero-shot generalization scenarios, significantly outperforming state-of-the-art end-to-end policy models such as VLAs and WAMs. Project website at worldactionplanner.github.io