School of Computer Science, The University of Sydney
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which trains VLA policies to handle missing visual inputs and temporarily supplies predicted images when cameras fail. It uses optical-flow extrapolation for missing wrist views or all-vision loss, and an action-conditioned world model for missing third-person views, withdrawing the predictions when they become unreliable. Experiments on π0.5 and GR00T N1.5 show that MINT improves task success under camera loss over the original models. Experiments on the AgiBot G2 further demonstrate real-robot closed-loop deployment under camera loss.
A paired, closed-loop benchmark for evaluating VLA models with missing camera inputs, built on RoboCasa365. Each policy is measured against its own successful rollouts, so the score isolates what camera loss takes away.
For task j, Hj is the healthy success rate over its 50 scenes and Mc,j the fraction of those healthy successes that survive interruption condition c. The task score averages Hj with the nine conditional scores, and the MAIL-Bench score averages the 18 task scores. A policy is therefore rewarded both for solving tasks and for keeping what it can already do when a camera disappears.
| Benchmark | Visual failure | Onset | Post-loss control |
|---|---|---|---|
| COLOSSEUM | Perturbation | – | – |
| RoboTwin 2.0 | Perturbation | – | – |
| VLATest | Perturbation | – | – |
| LIBERO-Plus | Black frame | Start | Closed loop |
| Missing-modality IL | Camera dropout | Start | Open loop |
| MAIL-Bench (ours) | Camera loss | Mid-episode | Closed loop |
MINT separates robustness to missing visual inputs from visual prediction. The policy is first trained to act under variable camera availability, then selectively receives predicted views at inference time when a camera is interrupted.
During fine-tuning, camera groups are dropped as hard missing: their tokens are excluded and the availability flag is set to 0. The policy learns to act from whatever views remain, and healthy performance does not degrade; on GR00T N1.5 it improves on two of the three task categories.
When the wrist view or all vision is lost, short-horizon optical-flow extrapolation of the last frames fills the gap on the CPU. When the third-person view is lost, an action-conditioned world model generates it from the surviving wrist stream and the executed actions.
Each prediction is checked against the views that remain; once a route becomes unreliable, its images are withdrawn and the policy returns to acting under hard missing, which the training stage prepared it for. Predictions are supplied temporarily, never trusted indefinitely.
MINT does not increase the per-step wall-clock time of healthy execution. Under interruption the optical-flow route runs on the CPU, and the world model adds 9.2 GiB of GPU memory and about 100 TFLOPs per generation, invoked only while a third-person view is missing.
Healthy is the success rate with all cameras available. W, T and B report the scores under wrist, third-person and all-vision interruption at 30 / 45 / 60% of completion time. AVG is the task-category MAIL-Bench score and Δ its change over the corresponding base checkpoint. Bold marks the best result within each backbone and task category.
| Method | Tasks | Healthy | Wrist missing | Third-person missing | All vision missing | AVG | Δ | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| W30 | W45 | W60 | T30 | T45 | T60 | B30 | B45 | B60 | |||||
| GR00T N1.5 | |||||||||||||
| base model | pick-and-place | 0.732 | 0.163 | 0.276 | 0.341 | 0.165 | 0.224 | 0.424 | 0.000 | 0.000 | 0.010 | 0.234 | – |
| small appliances | 0.290 | 0.153 | 0.277 | 0.398 | 0.245 | 0.296 | 0.403 | 0.110 | 0.170 | 0.226 | 0.257 | – | |
| fixtures & navigation | 0.386 | 0.178 | 0.253 | 0.385 | 0.308 | 0.319 | 0.398 | 0.029 | 0.084 | 0.119 | 0.246 | – | |
| + MINT | pick-and-place | 0.724 | 0.423 | 0.561 | 0.673 | 0.622 | 0.674 | 0.647 | 0.053 | 0.091 | 0.166 | 0.463 | +0.229 |
| small appliances | 0.450 | 0.493 | 0.536 | 0.631 | 0.602 | 0.628 | 0.573 | 0.307 | 0.419 | 0.528 | 0.517 | +0.260 | |
| fixtures & navigation | 0.497 | 0.450 | 0.507 | 0.649 | 0.335 | 0.488 | 0.566 | 0.076 | 0.197 | 0.378 | 0.414 | +0.168 | |
| π0.5 | |||||||||||||
| base model | pick-and-place | 0.636 | 0.059 | 0.098 | 0.200 | 0.239 | 0.245 | 0.375 | 0.000 | 0.014 | 0.089 | 0.196 | – |
| small appliances | 0.257 | 0.270 | 0.264 | 0.342 | 0.291 | 0.343 | 0.413 | 0.065 | 0.095 | 0.156 | 0.250 | – | |
| fixtures & navigation | 0.360 | 0.151 | 0.266 | 0.355 | 0.453 | 0.469 | 0.602 | 0.043 | 0.129 | 0.254 | 0.308 | – | |
| + MINT | pick-and-place | 0.688 | 0.027 | 0.111 | 0.342 | 0.574 | 0.735 | 0.713 | 0.025 | 0.090 | 0.196 | 0.350 | +0.154 |
| small appliances | 0.333 | 0.291 | 0.426 | 0.350 | 0.513 | 0.372 | 0.480 | 0.218 | 0.337 | 0.370 | 0.369 | +0.119 | |
| fixtures & navigation | 0.414 | 0.131 | 0.261 | 0.528 | 0.428 | 0.498 | 0.546 | 0.093 | 0.128 | 0.278 | 0.331 | +0.023 | |
Per-task, per-condition results and bootstrap intervals are in the paper's appendix.
The same observation interface transfers to a physical robot without changing the VLA architecture. An availability-aware GR00T N1.5 policy is fine-tuned on 203 successful teleoperated episodes of four tabletop tasks; the G2 head camera provides the third-person view and the right wrist camera the wrist view.
Healthy episodes; head (third-person) camera on top and right wrist camera below, at five time points.




@article{jiang2026mint,
title = {Learning to Act under Visual Interruptions with Vision-Language-Action Models},
author = {Jiang, Mingle and Xu, Rui and Wang, Yunke and Xu, Chang},
journal = {arXiv preprint},
year = {2026}
}