Learning to Act under Visual Interruptions with Vision-Language-Action Models

Mingle Jiang, Rui Xu, Yunke Wang, Chang Xu

School of Computer Science, The University of Sydney

Head camera interrupted, real time. Left: the real head camera, hidden from the policy after 6.2 s. Middle: the head view as the policy receives it: real frames until the interruption, then the world-model prediction MINT supplies at each policy query (blue), and no frame where none is accepted. Right: the wrist camera, which stays available. GR00T N1.5 + MINT completes the task: cube into the drawer, drawer closed.
MAIL-Bench: four availability conditions on RoboCasa365 and the number of scenes solved by pi0.5 and GR00T N1.5 with every camera present and after one visual role is lost at 30, 45 or 60 percent of completion time, with and without MINT
MAIL-Bench. (a) Four availability conditions on RoboCasa365; an interrupted camera delivers no frame and its availability flag is 0. (b) Scenes solved out of 900 with every camera present (healthy) and after one visual role is lost at 30, 45 or 60% of each policy's own completion time. Camera loss removes a large share of both base policies' successes, and MINT recovers part of that loss on every role.

Abstract

Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which trains VLA policies to handle missing visual inputs and temporarily supplies predicted images when cameras fail. It uses optical-flow extrapolation for missing wrist views or all-vision loss, and an action-conditioned world model for missing third-person views, withdrawing the predictions when they become unreliable. Experiments on π0.5 and GR00T N1.5 show that MINT improves task success under camera loss over the original models. Experiments on the AgiBot G2 further demonstrate real-robot closed-loop deployment under camera loss.

MAIL-Bench: manipulation under visual interruption

A paired, closed-loop benchmark for evaluating VLA models with missing camera inputs, built on RoboCasa365. Each policy is measured against its own successful rollouts, so the score isolates what camera loss takes away.

Benchmark construction: a healthy rollout defines the reference completion time, then the same scene is replayed with the selected visual role interrupted at 30, 45 or 60 percent of that time and kept unavailable
A successful healthy rollout defines the reference completion time Thealthy. The same scene is then replayed with the selected visual role interrupted at 30%, 45% or 60% of Thealthy and kept unavailable thereafter, while unaffected cameras continue streaming. Shown is an AgiBot G2 drawer example.

Protocol

  • Environment. The 18 RoboCasa365 Atomic-Seen tasks with 50 fixed scenes each: 900 evaluation scenes.
  • Visual roles, not devices. The wrist camera forms the wrist group; the base-mounted cameras form the third-person group. Conditions: healthy, wrist missing, third-person missing, all vision missing.
  • Hard missing. Once interrupted, a camera group provides no further frames and its availability flag is set to 0. No black frames, no frozen images.
  • Onset relative to the policy. Interruptions start at 30%, 45% and 60% of each policy's own healthy completion time, so fast and slow policies are compared at the same task stage.
  • Paired and deterministic. Every interruption rollout restarts from the healthy rollout's saved initial state and policy seed; the pre-interruption trajectory is checked to match the reference.

What it measures

For task j, Hj is the healthy success rate over its 50 scenes and Mc,j the fraction of those healthy successes that survive interruption condition c. The task score averages Hj with the nine conditional scores, and the MAIL-Bench score averages the 18 task scores. A policy is therefore rewarded both for solving tasks and for keeping what it can already do when a camera disappears.

How it differs from robustness benchmarks

BenchmarkVisual failureOnsetPost-loss control
COLOSSEUMPerturbation––
RoboTwin 2.0Perturbation––
VLATestPerturbation––
LIBERO-PlusBlack frameStartClosed loop
Missing-modality ILCamera dropoutStartOpen loop
MAIL-Bench (ours)Camera lossMid-episodeClosed loop

MINT: acting under visual interruption

MINT separates robustness to missing visual inputs from visual prediction. The policy is first trained to act under variable camera availability, then selectively receives predicted views at inference time when a camera is interrupted.

MINT overview: availability-aware training, then at inference optical-flow extrapolation for missing wrist views or all-vision loss and an action-conditioned world model for missing third-person views, with reliability-gated fallback
MINT overview. Unreliable predictions are withdrawn, and the policy falls back to acting from the observations that remain.

TrainLearning with missing visual inputs

During fine-tuning, camera groups are dropped as hard missing: their tokens are excluded and the availability flag is set to 0. The policy learns to act from whatever views remain, and healthy performance does not degrade; on GR00T N1.5 it improves on two of the three task categories.

FlowWorldPredicting missing visual inputs

When the wrist view or all vision is lost, short-horizon optical-flow extrapolation of the last frames fills the gap on the CPU. When the third-person view is lost, an action-conditioned world model generates it from the surviving wrist stream and the executed actions.

GateReliability-gated fallback

Each prediction is checked against the views that remain; once a route becomes unreliable, its images are withdrawn and the policy returns to acting under hard missing, which the training stage prepared it for. Predictions are supplied temporarily, never trusted indefinitely.

Cost

MINT does not increase the per-step wall-clock time of healthy execution. Under interruption the optical-flow route runs on the CPU, and the world model adds 9.2 GiB of GPU memory and about 100 TFLOPs per generation, invoked only while a third-person view is missing.

Results on MAIL-Bench

Healthy is the success rate with all cameras available. W, T and B report the scores under wrist, third-person and all-vision interruption at 30 / 45 / 60% of completion time. AVG is the task-category MAIL-Bench score and Δ its change over the corresponding base checkpoint. Bold marks the best result within each backbone and task category.

MethodTasksHealthyWrist missingThird-person missingAll vision missingAVGΔ
W30W45W60T30T45T60B30B45B60
GR00T N1.5
base modelpick-and-place0.7320.1630.2760.3410.1650.2240.4240.0000.0000.0100.234–
small appliances0.2900.1530.2770.3980.2450.2960.4030.1100.1700.2260.257–
fixtures & navigation0.3860.1780.2530.3850.3080.3190.3980.0290.0840.1190.246–
+ MINTpick-and-place0.7240.4230.5610.6730.6220.6740.6470.0530.0910.1660.463+0.229
small appliances0.4500.4930.5360.6310.6020.6280.5730.3070.4190.5280.517+0.260
fixtures & navigation0.4970.4500.5070.6490.3350.4880.5660.0760.1970.3780.414+0.168
π0.5
base modelpick-and-place0.6360.0590.0980.2000.2390.2450.3750.0000.0140.0890.196–
small appliances0.2570.2700.2640.3420.2910.3430.4130.0650.0950.1560.250–
fixtures & navigation0.3600.1510.2660.3550.4530.4690.6020.0430.1290.2540.308–
+ MINTpick-and-place0.6880.0270.1110.3420.5740.7350.7130.0250.0900.1960.350+0.154
small appliances0.3330.2910.4260.3500.5130.3720.4800.2180.3370.3700.369+0.119
fixtures & navigation0.4140.1310.2610.5280.4280.4980.5460.0930.1280.2780.331+0.023

Per-task, per-condition results and bootstrap intervals are in the paper's appendix.

What the numbers say

  • Availability-aware training alone lifts healthy success on GR00T N1.5 (small appliances 0.290 → 0.450): learning under partial observation does not trade off normal execution.
  • Under interruption MINT outperforms both base models in almost all of the nine conditions, across wrist, third-person and all-vision loss and across early and late onsets.
  • Third-person loss benefits most from prediction, because the global view can still be inferred from the wrist camera; all-vision loss is the hardest condition for every policy.
Ablation on GR00T N1.5: MAIL-Bench scores at each onset for the base model, missing-view training only, prediction only, and full MINT
Ablation on GR00T N1.5. Scores averaged over the three task categories at each onset; H marks the healthy-input score of each variant. Missing-view training governs how far the policy falls when a view is lost early; prediction determines how much of a lost view can be recovered while another view remains (+0.105 at T30).

Real robot: AgiBot G2

The same observation interface transfers to a physical robot without changing the VLA architecture. An availability-aware GR00T N1.5 policy is fine-tuned on 203 successful teleoperated episodes of four tabletop tasks; the G2 head camera provides the third-person view and the right wrist camera the wrist view.

Success rate of GR00T N1.5 with and without MINT on the AgiBot G2 under healthy execution and under sustained third-person, wrist and all-camera interruption, 10 trials per task and condition
Real robot experiment on the AgiBot G2. Success rate of GR00T N1.5 and MINT over 10 trials per task and condition; numbers above the bars are successful trials.
A cube-to-drawer episode on the AgiBot G2 with the third-person camera interrupted at step 50: real third-person stream hidden from the policy, the world-model prediction that replaces it, and the wrist stream that stays available
Episode on the cube-to-drawer task with the third-person camera interrupted at step 50. Top: real third-person stream, hidden from the policy after interruption. Middle: the third-person stream and its world-model prediction. Bottom: the wrist stream, which stays available. The policy completes grasping, transport, release and drawer closure.

The four tasks

Healthy episodes; head (third-person) camera on top and right wrist camera below, at five time points.

Healthy drawer episode
(a) Drawer
Healthy cube-into-bowl episode
(b) Cube into bowl
Healthy cup-onto-plate episode
(c) Cup onto plate
Healthy stack-cups episode
(d) Stack cups

BibTeX

@article{jiang2026mint,
  title   = {Learning to Act under Visual Interruptions with Vision-Language-Action Models},
  author  = {Jiang, Mingle and Xu, Rui and Wang, Yunke and Xu, Chang},
  journal = {arXiv preprint},
  year    = {2026}
}