Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which first trains VLA policies to remain functional under missing visual inputs. At inference time, MINT selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable. Experiments on π0.5 and GR00T N1.5 show that MINT significantly improves task success under camera loss over the original models. Experiments on AgiBot G2 further demonstrate the real-robot deployment under camera loss. The benchmark is available at https://minglejiang.github.io/Mail-Bench/
Figures & tables
Figure 1: MAIL-Bench. (a) Four availability conditions on RoboCasa365; an interrupted camera delivers no frame and its availability flag is 0. (b) Scenes solved out of 900 with every camera present (healthy) and after one visual role is lost at 30, 45 or 60% of each policy’s own completion time: camera loss removes a large share of both base policies’ successes, and MINT recovers part of that loss on every role.
Benchmark
Visual failure
Onset
Post-loss control
COLOSSEUM ( Pumacay et al., 2024 )
Perturbation
–
–
RoboTwin 2.0 ( Chen et al., 2026 )
Perturbation
–
–
VLATest ( Wang et al., 2025 )
Perturbation
–
–
LIBERO-Plus ( Fei et al., 2026 )
Black frame
Start
Closed loop
Missing-modality IL ( Ismkhan and Bouchahcia, 2026 )
Camera dropout
Start
Open loop
MAIL-Bench (ours)
Camera loss
Mid-episode
Closed loop
Table 1: Comparison of visual-robustness evaluation settings. MAIL-Bench uniquely targets persistent visual-stream loss during closed-loop execution.
Figure 2: A successful healthy rollout defines the reference completion time Thealthy . The same scene is then replayed with the selected visual role interrupted at 30%, 45%, or 60% of Thealthy and kept unavailable thereafter, while unaffected cameras continue streaming. Shown is an AgiBot G2 drawer example.
Figure 3: MINT overview. MINT separates robustness to missing visual inputs from visual prediction. The policy is first trained to act under variable camera availability, then selectively receives predicted views at inference time when a camera is interrupted. Unreliable predictions are withdrawn, and the policy falls back to acting from the observations that remain.
Method
Tasks
Healthy
Wrist missing
Third-person missing
All vision missing
AVG
Δ
W30
W45
W60
T30
T45
T60
B30
B45
B60
GR00T N1.5 ( NVIDIA et al., 2025 )
base model
pick-and-place
0.732
0.163
0.276
0.341
0.165
0.224
0.424
0.000
0.000
0.010
0.234
-
small appliances
0.290
0.153
0.277
0.398
0.245
0.296
0.403
0.110
0.170
0.226
0.257
-
fixtures & navigation
0.386
0.178
0.253
0.385
0.308
0.319
0.398
0.029
0.084
0.119
0.246
-
+MINT
pick-and-place
0.724
0.423
0.561
0.673
0.622
0.674
0.647
0.053
0.091
0.166
0.463
+0.229
Table 2: MAIL-Bench results. Healthy denotes success rate with all cameras available, while W/T/B report scores under wrist, third-person, and all-vision interruption at 30/45/60% of completion time. AVG is the MAIL-Bench score and Δ its improvement over the corresponding base checkpoint. Bold marks the best result within each backbone and task category.
Figure 4: Ablation on GR00T N1.5. Scores averaged over the three task categories at each interruption onset. “H” marks the healthy-input score of each variant.
Figure 5: Episode on ‘Cube to Drawer’ task with the third-person camera interrupted at step 50. The policy completes grasping, transport, release and drawer closure. Top: real third-person stream, hidden from the policy after interruption. Middle: the third-person view stream and prediction by the world model. Bottom: the wrist stream, which stays available.
Figure 6: Real robot experiment on the AgiBot G2. Success rate of GR00T N1.5 and MINT over 10 trials per task and condition. Numbers above the bars are successful trials.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Horizon
Task
Horizon
CloseBlenderLid
900
PickPlaceCounterToStove
600
CloseFridge
900
PickPlaceDrawerToCounter
750
CloseToasterOvenDoor
450
PickPlaceSinkToCounter
900
CoffeeSetupMug
600
PickPlaceToasterToCounter
600
NavigateKitchen
450
SlideDishwasherRack
450
OpenCabinet
1050
TurnOffStove
750
Appendix
Table 3: Task suite. The 18 RoboCasa365 Atomic-Seen tasks and their horizons in control steps (20 Hz).
Hyperparameter
Value
Optical-flow route
Min. flow confidence (wrist / third-person)
0.05 / 0.08
Motion budget, wrist (end-effector + base)
0.15 m / 0.70 rad
Motion budget, third-person (base)
0.30 m / 0.70 rad
Minimum tenancy
64 steps
World-model route
Appendix
Table 4: Hyperparameters of MINT.
Wrist missing
Third-person missing
All vision missing
Policy
H
W30
W45
W60
T30
T45
T60
B30
B45
B60
SMAIL
π0.5
0.402
0.165
0.219
0.307
0.340
0.365
0.476
0.038
0.086
0.176
0.257
π0.5 + MINT
0.463
0.156
0.274
0.417
0.497
0.522
0.571
0.116
0.187
0.286
0.349
GR00T N1.5
0.450
0.165
0.267
0.377
0.247
0.285
0.407
0.048
0.089
0.124
0.246
GR00T N1.5 + MINT
0.544
0.457
0.532
0.650
0.504
0.586
0.591
0.147
0.242
0.369
0.462
Appendix
Table 5: Main results. Healthy success H , conditional scores Mc under wrist (W), third-person (T) and all-vision (B) interruption at the 30/45/60% onsets, and SMAIL (Eq. 2 ), averaged over the 18 tasks.
Wrist missing
Third-person missing
All vision missing
Arm
H
W30
W45
W60
T30
T45
T60
B30
B45
B60
SMAIL
Δ
Base
0.450
0.165
0.267
0.377
0.247
0.285
0.407
0.048
0.089
0.124
0.246
Clean FT + prediction
0.532
0.294
0.369
0.580
0.316
0.419
0.455
0.079
0.191
0.279
0.351
+0.105
Missing-view FT
0.540
0.412
0.542
0.643
0.405
0.495
0.605
0.151
0.252
0.334
0.438
+0.086
MINT
0.544
0.457
0.532
0.650
0.504
0.586
0.591
0.147
0.242
0.369
0.462
+0.024
Appendix
Table 6: GR00T N1.5 ablation. All variants except the base are served at the same eight-step cadence. Clean FT + prediction is fine-tuned on the same data and steps as MINT without missing-view examples; Missing-view FT is the MINT checkpoint without generated views. Δ is SMAIL minus the row above.
Arm
SMAIL
95% interval
Paired difference
ΔSMAIL
95% interval
π0.5
0.257
[0.221, 0.275]
π0.5 + MINT −π0.5
+0.091
[+0.063,+0.129]
π0.5 + MINT
0.349
[0.325, 0.365]
GR00T N1.5 + MINT − GR00T N1.5
+0.216
[+0.190,+0.240]
GR00T N1.5
0.246
[0.227, 0.265]
Clean FT + prediction − GR00T N1.5
+0.105
[+0.079,+0.130]
Clean FT + prediction
0.351
[0.332, 0.369]
Missing-view FT − Clean FT + prediction
+0.086
[+0.061,+0.109]
Missing-view FT
0.438
[0.417, 0.455]
GR00T N1.5 + MINT − Missing-view FT
+0.024
[+0.008,+0.044]
GR00T N1.5 + MINT
0.462
[0.443, 0.479]
GR00T N1.5 + MINT − Clean FT + prediction
+0.111
[+0.087,+0.133]
Appendix
Table 7: Bootstrap 95% intervals. Scenes are resampled within each task (10,000 resamples), and differences between arms are paired on the same resampled scenes. Top: SMAIL and paired differences. Bottom: per-condition paired differences between each backbone with MINT and its base.
Task
NjH
W30
W45
W60
T30
T45
T60
B30
B45
B60
Sj
CloseBlenderLid
0
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
0.000
CloseFridge
23
0.70
0.57
0.74
0.09
0.17
0.39
0.09
0.26
0.57
0.403
CloseToasterOvenDoor
2
0.00
1.00
1.00
0.50
0.50
1.00
0.00
0.50
1.00
0.554
CoffeeSetupMug
5
0.00
0.00
0.00
0.20
0.20
0.20
0.00
0.00
0.00
0.070
NavigateKitchen
0
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
0.000
OpenCabinet
34
0.03
0.12
0.18
0.09
0.15
0.15
0.00
0.00
0.00
0.139
Appendix
Table 8: Per-task, per-condition results for π0.5 .
Task
NjH
W30
W45
W60
T30
T45
T60
B30
B45
B60
Sj
CloseBlenderLid
2
0.00
0.50
0.00
0.50
0.00
0.00
0.00
0.00
0.00
0.104
CloseFridge
23
0.70
0.74
0.78
0.48
0.48
0.70
0.26
0.22
0.70
0.550
CloseToasterOvenDoor
9
0.11
0.22
0.56
0.11
0.22
0.33
0.22
0.22
0.33
0.251
CoffeeSetupMug
13
0.00
0.00
0.00
0.46
0.38
0.69
0.00
0.00
0.15
0.195
NavigateKitchen
1
0.00
0.00
1.00
0.00
0.00
0.00
0.00
0.00
0.00
0.102
OpenCabinet
36
0.06
0.22
0.31
0.11
0.36
0.25
0.00
0.03
0.03
0.208
Appendix
Table 9: Per-task, per-condition results for π0.5 + MINT.
Task
NjH
W30
W45
W60
T30
T45
T60
B30
B45
B60
Sj
CloseBlenderLid
3
0.00
0.00
0.67
0.00
0.00
0.00
0.00
0.00
0.00
0.073
CloseFridge
35
0.57
0.71
0.74
0.17
0.31
0.37
0.20
0.29
0.46
0.453
CloseToasterOvenDoor
12
0.25
0.50
0.92
0.08
0.17
0.42
0.00
0.00
0.08
0.266
CoffeeSetupMug
5
0.00
0.00
0.00
0.00
0.00
0.20
0.00
0.00
0.00
0.030
NavigateKitchen
5
0.20
0.20
0.20
0.40
0.00
0.00
0.00
0.20
0.00
0.130
OpenCabinet
32
0.19
0.25
0.44
0.06
0.06
0.16
0.00
0.03
0.00
0.183
Appendix
Table 10: Per-task, per-condition results for GR00T N1.5.
Task
NjH
W30
W45
W60
T30
T45
T60
B30
B45
B60
Sj
CloseBlenderLid
2
0.00
0.00
0.00
0.50
0.50
0.00
0.00
0.00
0.00
0.104
CloseFridge
30
0.83
0.70
0.80
0.30
0.53
0.63
0.23
0.20
0.37
0.520
CloseToasterOvenDoor
19
0.26
0.37
0.63
0.37
0.26
0.37
0.00
0.05
0.32
0.301
CoffeeSetupMug
11
0.18
0.18
0.64
0.36
0.36
0.64
0.09
0.00
0.09
0.277
NavigateKitchen
14
0.43
0.57
0.50
0.00
0.07
0.14
0.00
0.00
0.07
0.207
OpenCabinet
41
0.78
0.73
0.78
0.22
0.39
0.54
0.05
0.07
0.15
0.453
Appendix
Table 11: Per-task, per-condition results for GR00T N1.5 + MINT.
Figure 7: Healthy episodes of the four AgiBot G2 tasks. (a) drawer, (b) cube into bowl, (c) cup onto plate, (d) stack cups. Each panel shows the head (third-person) camera (top) and the right wrist camera (bottom) at five time points.