Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how π0.5 and GR00T models act under input faults such as image blackouts and freezing. We find that blackout and freezing produce distinct physical failure modes even when task-success rates are similarly low: freezing causes more extreme joint behavior, whereas blackout after gripper closure can cause more object drops, most markedly without proprioception. Selective intervention studies reveal that proprioception (current robot state) partly compensates for the removed robot depictions and reduces non-target contact. However, it cannot sufficiently restore task success when wrist-view object information is removed, even when aided by the remaining scene view. We then evaluate two mitigation approaches: camera-blackout training and training-free replacement of faulty visual embeddings. Both improve task success in selected conditions, but can increase unintended contact or disturbance to surrounding objects. Real-robot trials further show that successful execution under camera faults can still involve unintended physical interactions. These findings motivate designing VLA policies that use the robot and object information still available under camera faults to limit hazardous motion.
Figures & tables
Figure 1: Overview of the camera-fault evaluation protocol.
Figure 2: Task performance and executed commands under camera faults, with and without proprioception. (a) Task success under clean input, blackout, and freezing, pooling 150 episodes from each of Spatial, Object, and Goal ( n=450 per condition). (b) Gripper commands under clean, scene-blackout, and wrist-blackout inputs. Each panel displays the same 30 selected episodes, ordered within a policy by the clean first-close command; success labels use all 150 Spatial episodes. Gray denotes open commands and white denotes time after termination.
Measure
Event or statistic
Reported as
Core physical outcomes
NTC
Robot–object contact, excluding target and arena
% of steps
JVE
At least one joint exceeds its velocity limit
% of episodes or steps
JLP
Normalized J2 or J7 position enters the registered 0.05 limit margin
% of episodes
OD
Non-target object displacement >2 cm
% of episodes
Additional physical measures
Table 1: Safety-relevant physical measures.
Execution context
Contact and object effects
Joint motion
Proprio.
Camera
Fault
SR
∥Δxyz∥
NTC
OD
Drop
F5
JVE
JLP
π0.5
Yes
–
Clean
95 .3 → 95 .3
0 .76 → 0 .76
2 .0 → 2 .1
2 .0 → 1 .3
1 .3
0 .0 → 0 .0
0 .0 → 0 .0
1 .3 → 1 .3
Scene
Black
23 .3 → 66 .0
0 .58 → 0 .67
7 .2 → 2 .6
14 .0 → 6 .7
2 .0
1 .3 → 2 .0
0 .0 → 0 .0
2 .0 → 6 .0
Frozen
2 .7 → 6 .7
0 .71 → 0 .70
11 .7 → 3 .4
39 .3 → 12 .7
5 .3
36 .7 → 0 .7
85 .3 → 41 .3
85 .3 → 42 .7
Wrist
Black
21 .3 → 77 .3
0 .67 → 0 .71
1 .4 → 1 .7
5 .3 → 2 .0
4 .0
0 .0 → 0 .0
0 .0 → 0 .0
4 .7 → 2 .7
Table 2: Physical indicators under camera faults. Paired entries report results in the order of fault introduction: episode start → after gripper closure, with corresponding clean references. NTC reports the percentage of steps; SR, OD, Drop, F5 , JVE, and JLP report percentages of episodes. Drop is reported only for faults introduced after gripper closure. ∥Δxyz∥ denotes mean translation-command magnitude in native action units. Metric definitions and measurement protocols appear in Appendix B .
Figure 3: Real-world outcomes on the WidowX block-in-bowl task, with ten trials per condition and faults introduced from episode start. Left: Task success; labels give successful trial counts and error bars show Wilson 95% confidence intervals. Right: Operator-annotated non-target contact and drop incidence, and the median minimum joint-limit margin across trials. Unlike simulation NTC , contact is measured per trial. Joint-limit margin measures the minimum distance to configured arm-joint limits in radians; smaller values indicate closer proximity.
Figure 4: Task and physical outcomes of selective content removal on Spatial (150 episodes per cell). (a,b) Task success after removal at episode start. (c) Top: NTC under wrist-object removal versus rendering control. Bottom: JLP after removing scene arm links alone or together with the palm. Both physical indicators use episode-start interventions and up to 105 executed steps. Yes/No and + P/ − P indicate policies with/without proprioception. Exact success values appear in Table 10 ; measurement protocols appear in Appendix B .
Figure 5: Effects of blackout training on π0.5 on Spatial, with 150 episodes per condition per run. (a) Whiskers show observed ranges across two training runs where available. Full results appear in Table 14 . (b) Physical indicators use the first 105 executed steps; NTC and JVE are step percentages, and OD is an episode percentage.
Figure 6: Mean-embedding replacement under episode-start camera faults on Spatial (150 episodes per arm). (a) Task success for π0.5 and GR00T. (b) Physical outcomes for π0.5 .
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
π0.5
GR00T
Initialization
Base π0.5 weights
GR00T-N1.7-3B base weights
Evaluated update
10,000
4,000
Effective batch size
32
640
Optimizer
AdamW
AdamW
Learning rate
Linear warmup to 5×10−5
10−4 peak, cosine schedule
Warmup
10,000 updates
200 updates (5% of training)
Appendix
Table 3: Training settings for the principal simulation policies. The with- and without-proprioception policies are trained separately within each model family. These settings do not describe the released or short-budget checkpoints used in the additional mitigation experiments.
Ten tasks × 15 initial states = 150 episodes per suite and condition
Appendix
Table 4: Closed-loop simulation evaluation settings. Episode limits apply before any separately requested post-completion recording. Physical measurements use the windows specified in Appendix B .
Model
Proprio.
Clean
Scene black
Wrist black
Scene frozen
Wrist frozen
π0.5
Yes
93.1
28.9
13.1
5.8
3.6
π0.5
No
94.4
9.1
1.8
0.2
0.0
GR00T
Yes
98.7
51.6
12.2
9.3
0.0
GR00T
No
97.1
34.4
0.0
5.1
0.0
Appendix
Table 5: Task success (%) for policies with and without proprioception. Both models average Spatial, Object, and Goal ( n=450 per condition).
Model
Proprio.
Condition
Drop ( n )
Grasp ( n )
Drop (%)
Drop / grasp (%)
Camera faults
π0.5
Yes
Clean
2
146
1.3
1.4
π0.5
Yes
Scene blackout
3
137
2.0
2.2
π0.5
Yes
Scene freezing
8
144
5.3
5.6
π0.5
Yes
Wrist blackout
6
135
4.0
4.4
π0.5
Yes
Wrist freezing
3
138
2.0
2.2
Appendix
Table 6: Drop counts and grasp opportunity under camera faults. Every row contains 150 episodes. The reported Drop rate uses all 150 episodes; the conditional rate uses only episodes with an established grasp. A grasp requires bilateral finger contact and target height more than 2 cm above its first recorded height for three consecutive steps. Camera faults are introduced after gripper closure, with matched clean references.
Model
Proprio.
Condition
Grasp
Prefix
Full post50
Zero post50
π0.5
Yes
Wrist blackout
135
132.0
117
0
π0.5
Yes
Wrist freezing
138
135.0
126
0
π0.5
No
Wrist blackout
130
220.0
145
1
π0.5
No
Wrist freezing
134
220.0
136
1
GR00T
Yes
Wrist blackout
142
183.5
125
0
GR00T
Yes
Wrist freezing
144
220.0
141
0
Appendix
Table 7: Measurement opportunity in the representative contrasts. Every row has 150 episodes. Prefix is the median number of recorded steps through first completion (full trace for failures). Full post50 counts episodes with all 50 observed steps after logged fault onset before completion; zero post50 counts episodes with no such exposure. Episodes with shorter exposure remain in the all-episode denominator. Grasp is the number satisfying the established-grasp rule before completion.
Table 8: Uncertainty and measurement-window sensitivity for representative physical contrasts. Each contrast uses 150 paired Spatial episodes. A and B name the two conditions in each block. Rates are percentages of episodes, and Δ=A−B is an absolute difference on that scale. Primary intervals are 95% task-cluster bootstrap intervals from 10,000 resamples of the ten tasks. The primary window ends at first task completion, or at episode end for failures. Additional columns give differences for the first 105 steps, the first 50 steps from logged fault onset (both capped at first completion), and the full recorded trace. Drop uses the same contact-conditioned rule in every column; both contact loss and the qualifying descent must lie inside the selected window.
Figure 7: π0.5 executed action coordinates under clean inputs, scene blackout, and wrist blackout. Rows follow the clean first-close order; gray marks time after episode termination. Color scales are shared across conditions and panels for each coordinate.
Figure 8: GR00T executed action coordinates under clean inputs, scene blackout, and wrist blackout. Rows follow the clean first-close order; gray marks time after episode termination. Color scales are shared across conditions and panels for each coordinate.
Policy
Proprio.
Scene black
Scene frozen
Wrist black
Wrist frozen
π0.5
Yes
.94→.27
1.00→.33
.99→.74
1.00→−.10
π0.5
No
.85→.15
1.00→.24
.97→.55
.99→.11
GR00T
Yes
.99→.46
1.00→.33
.98→.76
.99→.07
GR00T
No
.99→.31
1.00→.12
.50→.06
.99→.13
Appendix
Table 9: Median xyz command cosine relative to paired clean execution. Each cell shows early [0,20)→ later [50,80) values over eligible episode-step pairs at ϵ=0.05 . A value near 1 indicates similar native command direction, irrespective of magnitude.
Figure 9: Immediate π0.5 gripper commands for 200 paired Spatial observations per phase, comparing clean inputs with both camera images blacked out. Commands are averaged over the planned five-step execution prefix. Open is on the left and close is on the right.
π0.5
GR00T
Proprioception
Yes
No
Yes
No
Scene view
Robot rendering control
95.3 (143)
94.0 (141)
97.3 (146)
94.0 (141)
Fingers hidden
95.3 (143)
96.0 (144)
96.0 (144)
96.7 (145)
Arm links hidden
93.3 (140)
90.7 (136)
96.7 (145)
88.7 (133)
Arm links and fingers hidden
89.3 (134)
75.3 (113)
92.7 (139)
81.3 (122)
Appendix
Table 10: Exact task success for the selective-removal conditions in Figure 4 . Cells show success percentage (successful episodes), with 150 LIBERO-Spatial episodes per condition. Robot rendering controls remove shadows; object rendering controls retain them. The other camera is unmodified.
Removed content
Policy
Proprio.
Success
Δxyz
NTC
OD
Scene robot
π0.5
Yes
.507
.034
.055
.133
π0.5
No
.133
.084
.082
.247
GR00T
Yes
.467
.049
.029
.073
GR00T
No
.233
.075
.045
.200
Wrist objects
π0.5
Yes
.027
.061
.095
.200
π0.5
No
.027
.110
.126
.327
Appendix
Table 11: Task, command, and physical outcomes after selective content removal. Translation reports RMS change from the matched rendering control over the first 20 executed steps. Physical indicators use up to 105 steps: NTC is non-target contact, and OD is sustained non-target object displacement. Each row contains 150 episodes.
Figure 10: Camera-disagreement inputs and policy responses. (a) Scene/wrist input pairs with matching views (outer columns) or different times or target object positions (middle columns). (b) Scene- and wrist-reference following across execution stages, averaged over both disagreement conditions. (c) Median translation-action projections toward the shifted-target reference, grouped by end-effector distance.
Model
Proprio.
Phase
n
Scene (%)
Wrist (%)
π0.5
Yes
Early episode
61
23.0
77.0
π0.5
Yes
Near close
100
41.5
58.5
π0.5
Yes
Transport
100
61.5
38.5
π0.5
Yes
Near episode end
65
27.7
72.3
π0.5
No
Early episode
61
30.3
69.7
π0.5
No
Near close
100
34.5
65.5
Appendix
Table 12: Camera-reference following at a 10-frame lag, pooled equally over the B/A and A/B input pairs in Figure 10(b) . Each of n selected observations contributes two queries. Shares use the closer full-seven-action reference; no exact ties occur. Within each policy and phase, selection retains reference separation at or above the median.
Model
Proprio.
Distance (cm)
n
Scene median
Wrist median
π0.5
Yes
[0,6)
1119
0.030
0.874
π0.5
Yes
[6,12)
569
0.115
0.772
π0.5
Yes
[12,24)
386
0.268
0.706
π0.5
Yes
[24,40)
166
0.377
0.613
π0.5
No
[0,6)
1125
0.020
0.948
π0.5
No
[6,12)
627
0.071
0.881
Appendix
Table 13: Median XYZ action projection toward the displaced-target reference by end-effector distance. Distance bins are in cm and use [lo,hi) ; n is selected probes per bin and is shared by scene and wrist within a row.
Recipe
Proprio.
nruns
Clean
Scene black
Scene frozen
None
Yes
1
93.3
22.0
2.7
None
No
1
93.3
14.0
1.3
Scene 10%
Yes
2
96.0 [95.3,96.7]
91.0 [89.3,92.7]
34.3 [30.0,38.7]
Scene 10%
No
1
92.7
86.0
14.7
Scene 20%
Yes
2
94.7 [94.0,95.3]
94.0 [93.3,94.7]
53.0 [45.3,60.7]
Scene 20%
No
1
95.3
91.3
38.0
Appendix
Table 14: Task success (%) for the evaluated π0.5 LIBERO-Spatial training recipe and proprioception condition. Entries are mean [minimum, maximum] across training runs; the dedicated nruns column gives the number of runs.
Suite
Proprio.
Training
Clean
Scene black
Scene frozen
Wrist black
Wrist frozen
Spatial
Yes
Plain
99.3
44.7
7.3
8.0
0.0
Spatial
Yes
Scene blackout
98.7
99.3
58.0
7.3
0.0
Spatial
Yes
Wrist blackout
98.7
34.7
0.0
93.3
0.0
Spatial
Yes
Scene + wrist
100.0
96.0
7.3
88.0
0.0
Spatial
No
Plain
97.3
32.7
0.7
0.0
0.0
Spatial
No
Scene blackout
99.3
97.3
6.7
0.0
0.0
Appendix
Table 15: GR00T task success at episode-start faults across Spatial, Goal, and Object (%; 150 rollouts per cell). Yes and No indicate policies with and without proprioception. Training uses 20% camera blackout. One training seed per recipe.
Transplanted component
Success (%)
None
22.7
Vision encoder + projector
38.0
VLM backbone
90.0
Action expert
26.0
All components (trained policy)
94.7
Appendix
Table 16: Component transplants from a scene-blackout-trained π0.5 policy with proprioception into the corresponding policy trained without blackout.
Proprio.
Camera
Fault
Policy/input
SR
NTC
JVE
OD
Yes
scene
black
Control / clean
0.953
0.025
0.000
0.013
Yes
scene
black
Control / fault
0.233
0.054
0.000
0.053
Yes
scene
black
Trained / fault
0.953
0.021
0.000
0.027
Yes
scene
frozen
Control / clean
0.953
0.025
0.000
0.013
Yes
scene
frozen
Control / fault
0.027
0.101
0.018
0.253
Yes
scene
frozen
Trained / fault
0.587
0.039
0.000
0.060
Appendix
Table 17: Task success and physical outcomes after blackout training on LIBERO-Spatial. SR is whole-episode success; contact and joint-speed exceedance are step rates, and displacement reports sustained movement of a non-target object. Physical measurements use up to 105 steps.
Figure 11: π0.5 task success and physical indicators after blackout training on Object. Open circles, squares, and dashed lines denote base-fault, camera-trained-fault, and base-clean results. Physical indicators use the first 105 executed steps. NTC and JVE report step percentages; OD reports episode percentages.
Figure 12: π0.5 task success and physical indicators after blackout training on Goal. Open circles, squares, and dashed lines denote base-fault, camera-trained-fault, and base-clean results. Physical indicators use the first 105 executed steps. NTC and JVE report step percentages; OD reports episode percentages. NTC and OD use the 120 episodes with a defined target out of 150 rollouts per condition.
Figure 13: π0.5 task success with proprioception under detector-triggered mean-embedding replacement. Open circles denote the unmodified policy and squares the replacement intervention. Each comparison uses 150 Spatial episodes paired within one evaluation setup.
Fault
Original
Replacement
Δ success (%) [95%]
McNemar p
Scene start · Black
34/150
97/150
+42.0 [+26.7, +58.7]
3.4e-15
Scene start · Frozen
9/150
102/150
+62.0 [+50.7, +73.3]
2.0e-28
Wrist start · Black
35/150
36/150
+0.7 [-10.0, +11.3]
1.000
Wrist start · Frozen
0/150
39/150
+26.0 [+14.7, +38.0]
3.6e-12
Scene Random · Black
67/150
111/150
+29.3 [+5.3, +52.7]
1.7e-07
Scene Random · Frozen
15/150
108/150
+62.0 [+48.0, +74.0]
6.0e-26
Appendix
Table 18: Task success without and with detector-triggered mean-embedding replacement for camera faults from episode start and random, grasp, and hold onsets. Start denotes episode start. Each paired LIBERO-Spatial condition has 150 episodes. Differences are absolute changes in success (%) with 95% task-cluster bootstrap intervals; p is exact McNemar.
Fault
Δ NTC (%) [95% CI]
Δ OD (%) [95% CI]
Scene start · Black
-0.5 [-2.6, +1.8]
+6.7 [+0.7, +13.3]
Scene start · Frozen
-5.0 [-7.5, -2.5]
-15.3 [-23.3, -7.3]
Wrist start · Black
+2.1 [+0.9, +3.5]
+2.0 [+0.0, +4.7]
Wrist start · Frozen
-8.6 [-11.3, -6.1]
-16.0 [-22.0, -10.0]
Scene Random · Black
+0.1 [-0.9, +1.2]
+2.0 [-2.0, +6.7]
Scene Random · Frozen
-4.9 [-7.1, -2.7]
-17.3 [-24.0, -10.7]
Appendix
Table 19: Changes in physical outcomes after detector-triggered mean-embedding replacement for the same paired conditions as Table 18 . Each cell reports replacement minus original with an episode-level 95% bootstrap interval, using up to 105 executed steps. Differences are absolute changes in NTC (% of steps) and OD (% of episodes). OD records non-target displacement above 2 cm for six consecutive samples.
Figure 14: Absolute input-group attention masses for scene vision, wrist vision, and state across clean and camera-fault probes. Values are attention masses multiplied by 100, measured over 96 fixed-probe episodes across four suites. This probe cohort is separate from the LIBERO-Spatial gain rollouts.
Input
State bias
Scene bias
Success (%)
Scene blackout
0
0
24.7
0.25
0
23.3
0.5
0
21.3
1
0
14.7
2
0
4.0
4
0
0.0
Appendix
Table 20: Task success under manually selected attention biases. Each row uses 150 Spatial episodes and the same π0.5 checkpoint with proprioception. Unlisted token groups have zero bias.
Policy or intervention
Scene blackout (%)
Wrist blackout (%)
Unmodified control
22.0
21.3
Full group adjustment
26.7
22.7
Camera-only adjustment
—
26.7
Blackout-trained reference
94.7
72.7
Appendix
Table 21: Attention biases calibrated to blackout-trained policies. Success uses 150 episodes per cell. Full adjustment targets state and camera groups; camera-only adjustment targets camera groups. The trained reference is a separate policy trained with blackout in the evaluated camera. A dash denotes a setting not shown in this summary.
Figure 15: Inference-time attention reweighting under matched scene- and wrist-blackout conditions. Each arm contains 150 Spatial episodes; intervals are 95% Wilson confidence intervals. Trained-checkpoint controls remain distinct from inference-only interventions.
Policy
Camera fault
No mask
Expert only
Prefix + expert
Released
Scene freezing
1.3
4.0
47.3
Wrist freezing
0.0
0.0
—
Fine-tuned, with proprio.
Scene freezing
25.3
30.0
18.0
Wrist freezing
0.0
0.0
—
Fine-tuned, without proprio.
Scene freezing
6.7
14.0
10.7
Wrist freezing
0.0
0.0
—
Appendix
Table 22: Task success (%) under oracle camera-token masking, with 150 episodes per cell. The two fine-tuned policies are a separate non-EMA checkpoint pair. All faults and masks start at episode start. Dashes indicate unevaluated settings.
Evaluation input
Control
Augmentation
Augmentation + freshness
Clean
42.0
36.0
46.0
Scene freezing
4.0
22.0
12.0
Wrist freezing
0.0
2.0
10.0
Appendix
Table 23: Task success (%) after 4,000 updates from the base checkpoint, with 50 episodes per cell. Freezing begins after three consecutive gripper-close commands. All scheduled episodes remain in the denominator, including those that do not reach the closing trigger. The control receives no temporal augmentation.
Temporal mixture
Evaluation input
Control
Delay 8–40
Delay 2–8
+ Consistency
Clean
98.7
98.7
98.7
73.3
Scene freezing
0.0
7.3
1.3
4.7
Wrist freezing
0.0
0.0
0.0
0.0
Wrist blackout
0.0
0.0
0.0
0.0
Appendix
Table 24: Task success (%) after 4,000 additional updates from the released LIBERO policy, with 150 episodes per cell. Both temporal mixtures contain landmark freezing and delayed images; only the sampled delay range differs. Consistency uses the 8–40-frame mixture. Faults begin at episode start.
Deploying Vision-Language-Action (VLA) models in real robotic systems requires robustness not only to semantic and perceptual variations, but also to embodiment-side faults that change how actions are physically realized. Real robots can experience joint-level changes caused by actuator degradation, hardware faults, safety limits, collision damage, or wear-induced friction. These faults are critical because they alter the action-to-motion interface of a policy, disrupting the learned closed-loop relationship between commanded actions, realized motion, and subsequent observations. In this work, we study realistic joint-level physical faults and show that VLA models are vulnerable when predicted actions are executed through a perturbed robot body. Our analysis reveals joint-dependent effects, with heterogeneous degradation in task success across affected joints. We also show that performance drops cannot be attributed solely to physical infeasibility, since feasible faults such as increased joint friction can still substantially reduce success rates and induce closed-loop execution mismatch. Motivated by these findings, we propose Joint-level Physical-fault Aware Residual Calibrator (J-PARC), a lightweight residual calibration framework built on top of a frozen VLA policy. J-PARC infers a latent joint-fault regime from recent joint dynamics and conditions a shared residual calibrator on this regime, enabling adaptive action correction across faulty joints. Experiments show that J-PARC improves robustness under joint-level faults while preserving fault-free environment performance.
Minsoo Jo, Taeju Kwon, Junha Chun +2
Graduate School of Data Science, Seoul National University
Vision-Language-Action (VLA) models demonstrate strong perfor-1 mance on language-conditioned robotic manipulation within their training dis-2 tribution, yet their generalization capabilities remain fundamentally limited. They3 lack the robustness required to handle perturbations, frequently failing when con-4 fronted with lighting changes, altered camera viewpoints, or small initial-state5 variations. We propose PROBEACT, a training-free runtime intervention frame-6 work that detects and recovers from grasping and placement failures in pre-7 trained VLA policies without modifying their weights or requiring additional8 demonstrations. PROBEACT combines three components: (i) a lightweight multi-9 target hidden-state probe that predicts the 3D positions of task-relevant objects10 from intermediate VLA features, with Hungarian-matched identity tracking for11 multi-object scenes; (ii) an object-agnostic kinematic state machine that detects12 grasp, transport, and placement failures using only gripper-internal signals and13 end-effector kinematics; and (iii) a hierarchical Control Barrier Function (CBF)14 filter that encodes repeated-failure locations as soft safe-set constraints, mini-15 mally correcting VLA actions while preserving baseline behavior. As a plug-and-16 play, training-free intervention loop, PROBEACT is orthogonal to existing train-17 ing pipelines. Evaluated on the LIBERO-plus benchmark, our framework acts as18 a universal safety net, improving the success rate of the OpenVLA-OFT model19 from 69.6% to 74.1%, while demonstrating broad applicability to both base and20 fine-tuned VLA policies.
Fan Zhang, Seongbin Park, Baharan Mirzasoleiman +2
University of California Los Angeles United States
Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions. We introduce ActFovea, a plug-and-play safeguarding framework that detects and mitigates such failures without retraining or modifying the underlying VLA policy. ActFovea uses robot kinematics, proprioceptive states, and recent actions to construct action-conditioned foveated regions that retain contact-relevant areas and predicted motion corridors while suppressing task-irrelevant visual content. It detects runtime risks by evaluating whether visual motion and observation freshness remain consistent with geometric, proprioceptive, and action transitions. For recoverable disturbances, ActFovea constructs disturbance-specific candidate observations and accepts a recovery only after verifying the resulting action chunk. When stale or replayed observations make reliable recovery impossible, it invokes a bounded safe-failure procedure. In closed-loop evaluations of π0 across multiple LIBERO suites, ActFovea increases success under localized visual overlays from 49.3% to 90.3%, closing 93.7% of the gap to clean performance. It further improves success under action drift and visual delay by 7.0 and 9.8 percentage points, respectively, while preserving clean-task performance. Under frozen-observation replay, ActFovea triggers timely safe failure in all trials, with no unprotected failures. These results demonstrate that spatiotemporal visual-action consistency provides an effective basis for runtime safeguarding of VLA policies.
Wenda Yu, Tianshi Wang, Fengling Li +3
Tongji University · Mohamed bin Zayed University of Artificial Intelligence · Shanghai Artificial Intelligence Laboratory +1