Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle in open-world conditions, as they are hand-tuned for specific scenarios and lack generalization. Vision-Language Models (VLMs) offer a promising alternative as they combine broad world knowledge with unified visual--text reasoning, enabling them to generalize across diverse scenarios and generate accurate, grounded task plans. However, for effective deployment in dynamic real-world settings, VLMs must be embedded into frameworks capable of handling uncertainty and environmental changes. Existing frameworks broadly address this reactively, triggering replanning only after execution failures or post-task checks, risking failed actions. Some methods verify conditions before actions, but these discrete checks miss changes occurring during execution. To address this, we present ProAct-VLM, an adaptive, physically grounded task planning framework that integrates VLMs within a real-time perception--feedback loop. ProAct-VLM continuously monitors the environment and re-plans as soon as relevant changes are detected, enabling adaptation before failure occurs. Evaluations against multiple baselines and across different VLM backbones show that our framework improves both success rates and efficiency in dynamic, long-horizon manipulation tasks. Project page: https://github.com/moured/ProAct-VLM
Figures & tables
Fig. 1: Overview of the proposed ProAct-VLM framework. (a) The system begins with a high-level user prompt, which is then converted into a planning prompt. (b) The VLM generates an initial plan based on this input. (e) Meanwhile,the Environment State Estimator continuously constructs a structured scene representation (object IDs, descriptions, poses, and the current Augmented image). (c) The Plan Monitor tracks this state and identifies task-relevant changes (e.g., object addition, removal, or goal modification). When such changes occur, (f) a re-planning prompt is issued to the VLM. (d) The Motion Planner executes the current plan until re-planning is triggered, enabling adaptive task execution in dynamic environments..
Fig. 2: Overview of the proposed perception pipeline. The system operates in three cyclic phases: Initialization (a), Tracking (b), and Environment Updating (c). In (a) Grounded-SAM detects objects from an input image using text prompts, and generates labeled segmentation masks. These masks are linked to CAD models and processed by FoundationPose for initial 6-DoF pose estimation. During (b), DeAOT tracks object masks across frames, while FoundationPose perform pose tracking for already registered objects. In (c), object detection is re-invoked periodically to identify new objects using a Comparing Mask Results (CMR) strategy, followed by pose estimation and reference update.
Fig. 3: illustrates of E in action. The framework detects, segments, and estimates object poses as the scene evolves, correctly handling object additions and removals.
Fig. 4: Disturbance types used in experiments: (a) object addition, (b) object removal, (c) goal change.
Fig. 5: Prompt templates used by VLM-Replan . The components in orange are specific to the LLM3+Image version, while the blue are specific to our Augmented version, the red text represents what we add to make the prompt for re-planning.
TABLE I: Defined disturbance scenarios used in our evaluation.
Scenario
GPT-4o
Gemini 2.5
Claude 4
Qwen 2.5
Llama 4
Succ
Eff
Succ
Eff
Succ
Eff
Succ
Eff
Succ
Eff
S1
100
100
100
100
100
100
20
100
100
100
S2
100
100
100
100
40
100
70
100
60
97
S3
80
80
100
100
0
0
60
70
100
100
S4
80
92.5
100
100
0
0
40
80
10
50
S5
100
96.6
100
100
0
0
50
90
100
96
TABLE II: Results of Ablation Study
Model
Prompt
S1
S2
S3
S4
S5
S6
S7
Avg
GPT-4o
LLM3+Image
90
40
80
30
100
60
30
61.4
OWG
100
80
70
70
100
90
70
83
Ours
100
100
80
80
100
90
70
89
VILA/REPLAN-VLM
100
0
100
40
80
30
0
50
Gemini 2.5 Pro
LLM3+Image
100
90
100
100
100
100
100
98.6
OWG
100
100
90
100
100
100
100
98.6
TABLE III: Results of Comparative Studies .
Fig. 6: The robot executes an initial VLM-generated plan; after a new object is introduced and detected by E (screenshot 8), the plan monitor triggers re-planning. The updated plan is executed successfully (screenshots 9–12).