Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known error, which requires producing a correction and judging whether it is right. Stopped at a failure and allowed to retry, they seldom repair it through their own randomness or from a language description of the error, and best-of-N selection cannot pick the successful candidate after a failure. We attribute this to training only on successful demonstrations and to inputs too narrow to show what went wrong, and conclude that reflection must come from a vision-language model (VLM), which takes in far more information, such as the episode history and text, and is more general. Prior VLM-led work has the VLM plan every step and invoke the embodied model as a tool, placing the VLM on the critical path. We propose Spotter, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control. We run Spotter with Qwen and with GPT as the VLM, and both improve the embodied models; with GPT, Spotter improves Cosmos Policy and π0.5 by 5.6 and 7.5 percentage points on RoboCasa, and raises π0.5 from 47.2% to 57.0% on the Hard setting of RoboTwin 2.0 and from 53% to 83% on a real robot. Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes only 13 to 16 s longer than with the embodied model alone and about 70% less time than with a VLM-led baseline using the same model. Our code is available at https://github.com/zqc3117/Spotter.
Figures & tables
Figure 1: The embodied model alone and with Spotter after the same missed grasp (task: move the lime from the pan to the plate). The grasp misses at step 160; by step 192 the arm heads for the plate. Alone (top), the model never notices, replaying the shape of a successful trajectory. With Spotter (bottom), the VLM notices the miss only after the arm has moved on: by step 192 it has already left the pan. This delay can be undone: the repair returns the arm to its pose before the miss and, since the failed attempt was off to the left, shifts the gripper right to face the lime.
Method
Architecture
Repair
Semantic
Training
Hi Robot
VLM-led
none
–
Yes
Harness VLA
VLM-led
failure memory
Partial
No
Towards the Harness
VLM-led
diagnose, retry
Yes
No
RACER
VLM-led
language hint
Yes
Yes
Spotter (ours)
Parallel
primitive plan
Yes
No
Table 1: Comparison with representative prior work. VLM-led designs wait on the external model; Spotter runs both in parallel. Semantic : corrects errors invisible in telemetry, such as reaching for the wrong object. Training : needs training beyond the off-the-shelf embodied model.
Figure 2: Error-correction ability of current embodied models. (a) Fraction of failed grasps that the model corrects when given three more attempts, before and after fine-tuning on recovery data, over 320 failed grasps per setting. (b) Pairwise ranking accuracy of test-time scorers and of the VLM judge given only images around the saved state, at normal states and after a failed grasp; error bars are 95% bootstrap confidence intervals.
Figure 3: Three ways to combine an embodied model with a VLM over time (block widths are schematic). (a) The embodied model acts alone. (b) VLM-led: the frontier VLM is called before every step, and the embodied model idles while it waits. (c) Spotter: a faster VLM screens each chunk in parallel, one chunk behind execution, and, on the routine path illustrated here, the frontier VLM is called after two consecutive flags. Once C2 and C3 are both flagged, the judge deliberates while the embodied model executes C5 ; it is interrupted only when the error is confirmed, at C6 , for the repair, then resumes. For clarity each drawn check covers one chunk and the lead limit K is not drawn; in practice a delayed check takes in every completed chunk not yet checked, so none is discarded.
Primitive
Effect
Move to point
Move the end effector to a back-projected target
Move by offset
Translate the end effector by a relative displacement
Lift
Raise the end effector
Rotate wrist
Rotate the wrist by a given angle
Set gripper
Open or close the gripper
Retreat
Return the arm to a recorded earlier pose of the judge’s choice
Table 2: Primitives available to the judge.
Short context
Full context
1-shot
Telemetry
last 3
last 40 (Cosmos) / 10 ( π0.5 )
as full
Turns
last 3
last 20
as full
Context restart
22k tokens
110k tokens
as full
Interventions per episode
3
8
8
Judge called after
1 flag
2 flags
2 flags
Example
none
none
1 episode
Table 3: The three Spotter configurations. Telemetry: recent chunks shown as rows; older chunks are summarized. Turns: recent Qwen judge turns kept in full, in addition to the initial briefing; older turns are compacted. Context restart: the token threshold for restarting with a text recap. GPT uses server-side conversation continuation.
Cosmos Policy
π0.5
Method
PnP
Δ
Non-PnP
Δ
All
Δ
PnP
Δ
Non-PnP
Δ
All
Δ
Embodied model
- 1.0× step limit
53.2
73.4
66.7
56.8
65.4
62.5
-Screener + fixed retry
52.2
74.2
66.9
56.2
65.0
62.1
-Oracle + fixed retry
53.5
73.9
67.1
56.0
66.6
63.1
- 1.8× step limit
55.2
74.2
67.9
58.0
67.5
64.3
Table 4: Success rates (%) on RoboCasa, 24 tasks with 50 episodes each (PnP: the 8 pick-and-place tasks; Non-PnP: the 16 tasks that operate a mechanism). Fixed retry : on a flagged failure the arm retreats and control returns to the embodied model, with no judge. Δ : gain over the embodied model at 1.8× the step limit, the budget of Spotter. Harness VLA uses 40 rounds and 5,000 steps (Appendix B ); with privileged info its VLM can also see the task’s success predicate, which Spotter never receives. Best per column in bold.
Figure 4: Success rate compared with other embodied models on (a) RoboCasa and (b) the Hard setting of RoboTwin 2.0. The compared methods are described in Appendix B .
Figure 5: Real-robot validation on a Franka Research 3. (a) The three tasks (Section 4.3 ). (b) Success rate over 10 trials per task for π0.5 alone and with Spotter (GPT), 1-shot.
Figure 6: Wall-clock time per episode on RoboCasa, overall and on successful and failed episodes: the embodied model alone at 1.8× the step limit, Spotter (Qwen, full context), and Harness VLA (Qwen, with privileged info), whose supervision blocks execution at every round. Time and tokens for every configuration are in Table 6 .
Judge calls
Repairs
Method
Succ.
Fail.
Succ.
Fail.
Cosmos Policy + Spotter (Qwen)
-short context
1.1
10.9
0.0
1.5
-full context
0.6
10.9
0.0
3.7
-1-shot
0.9
9.3
0.1
3.1
Cosmos Policy + Spotter (GPT)
Table 5: Judge calls and executed repairs per episode for every configuration, on successful and failed episodes. The π0.5 short-context row screens every two chunks, or 50 steps. For Harness VLA every interaction round counts as a repair.
Figure 7: The judge on RoboCasa, with Cosmos Policy and π0.5 pooled. Detection: share of the episodes it intervened in that the embodied model alone fails on the same scene. Correction: share of these errors corrected by a single repair, not by three attempts as in Section 3.1 .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Per episode
Successful episodes
Failed episodes
Method
Async
Sync
Tokens
Async
Sync
Tokens
Async
Sync
Tokens
(s)
(s)
(k)
(s)
(s)
(k)
(s)
(s)
(k)
Cosmos Policy
63
–
–
34
–
–
125
–
–
Spotter (Qwen)
-short context
158
196
82
54
77
18
395
467
227
-full context
181
217
87
47
71
9
521
587
283
Appendix
Table 6: Time and tokens per episode, overall and on successful and failed episodes. Async : wall-clock time with parallel supervision; Sync : time if supervision blocked execution, as in Harness VLA (with privileged info) at every round. Cosmos Policy and π0.5 : the models alone at 1.8× the step limit. Details in Appendix B .
Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and execution feedback. We introduce MachEmbodied-VLM (ME-VLM), a unified vision-language model with two variants, 4B and 35B-A3B, that brings together embodied cognition and multimodal agent capabilities. Our work emphasizes physical perception and spatiotemporal reasoning, together with planning, interaction, and outcome assessment in both digital and physical environments. We construct training data spanning embodied and multimodal agent tasks, including execution observations and feedback to support outcome assessment and decision refinement. The training pipeline comprises embodied capability injection, separate reinforcement learning of embodied and multimodal-agent experts, and multi-teacher on-policy distillation that consolidates their complementary capabilities into a single model. Experiments show competitive performance on both embodied and agent benchmarks, as well as on autonomous-driving and embodied-navigation tasks. For edge deployment, visual token compression, W4A8 quantization, and hardware-software co-optimization enable on-device inference of the 4B variant on the M100, reducing prefill latency from 400 ms to 188 ms. Project Page: https://machembodied.com/ME-Brain/ME-VLM.html Code Repository: https://github.com/MachEmbodied/ME-VLM
Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations into action plans. While prior work has studied abstention in LLMs, existing benchmarks are largely text-only and do not capture the perceptual grounding and physical constraints inherent to embodied robotics environments. In such settings, abstention requires recognizing when instructions are ambiguous, physically infeasible, based on false premises, or otherwise unresolvable given the available sensory modalities and context. To address this gap, we introduce a taxonomy to categorize abstention in the context of embodied robotics and present RoboAbstention, a scalable and auditable framework for generating abstention instructions grounded in images gathered from five robotics datasets. RoboAbstention instantiates the taxonomy through a three-phase pipeline: (1) structured visual grounding, (2) deterministic constraint derivation, and (3) controlled instruction generation via category-specific templates. This enables the construction of a diverse dataset with verifiable abstention conditions. We evaluate several frontier VLMs and find that all models exhibit significant weaknesses in abstention, including those with advanced reasoning capabilities. The best-performing model, Gemini 2.5 Flash, abstains on only 39.0% of our 6,069 benchmark instructions, while the embodied planner Gemini Robotics ER 1.6 Preview abstains on just 16.5%. We further explore methods for improving abstention in VLM planners, such as defensive prompting and in-context learning, and find that these interventions substantially improve performance, reaching 93.6% abstention rate for Gemini Robotics ER 1.6 Preview and 88.6% for GPT 5.4 Mini, yet no approach fully solves the problem. We open-source RoboAbstention at https://purseclab.github.io/RoboAbstention/.
Doguhan Yeke, Elif Su Temirel, Ananth Shreekumar +3
Purdue University · Bilkent University · Work performed while an intern at Purdue University.
In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visual cues. Existing VLA models learn a direct "Sense-to-Act" mapping from multimodal observations to robot actions. While effective within the training distribution, such tightly coupled policies are brittle under out-of-domain (OOD) shifts and difficult to correct when failures occur. Although recent embodied Chain-of-Thought (CoT) approaches expose intermediate reasoning, they still lack a mechanism for incorporating human spatial guidance, limiting their ability to resolve visual ambiguities or recover from mistakes. To address this gap, our framework allows users to optionally guide the policy with spatial priors, such as affordance points, boxes, and traces, which the subsequent reasoning process can directly condition on. Based on these inputs, the model generates a unified spatial-visual Chain-of-Thought that integrates external guidance with internal task planning, aligning human visual intent with autonomous decision-making. For practical deployment, we further couple the reasoning module with a lightweight reactive action head for efficient action execution. Extensive experiments demonstrate the effectiveness of our approach. On the in-domain SimplerEnv WidowX benchmark, our framework achieves a state-of-the-art 81.2% success rate. Under OOD visual shifts and spatial ambiguities, a single visual interaction substantially improves task success over existing methods, highlighting the value of interactive reasoning for failure recovery in embodied control. More details of the project can be found here: https://github.com/FutianLabs/GTA-VLA.
Yiran Ling, Qing Lian, Jinghang Li +6
Faculty of Computing, Harbin Institute of Technology · International Digital Economy Academy (IDEA) · National Key Laboratory of Smart Farm Technologies and Systems +4