Organizations: Fujitsu Limited, Kanagawa 211-8588, Japan. · Faculty of Science and Engineering, Waseda University, Tokyo 169-8050, Japan. · National Institute of Advanced Industrial Science and Technology, Tokyo 100-8921, Japan.
Imitation learning policies that integrate multiple sensory modalities are prone to overreliance on a dominant modality, such as vision, during training, which can disrupt policy execution when that modality is lost at inference time. In this paper, we introduce Targeted Modality Dropout (TMD), in which the dependence on each modality is estimated using attention and the most dominant modality is selectively dropped. This is combined with entropy regularization over the dependence distribution. Through real-robot evaluation using a bimanual manipulator, we show that under vision loss the success rate of the baseline policy drops substantially, whereas TMD sustains task execution. In contrast, a conventional dropout that selects the dropped modality at random, without the entropy regularization, fails on many tasks even without vision loss.
Figures & tables
Fig. 1: Overview of the problem and of the proposed method. Three modalities feed one policy. (a) Under ordinary training the dependence, estimated from attention, concentrates on vision. (b) TMD drops whichever modality is dominant during training, and the dependence balances. The bars are the estimated dependence at the end of training, with the vision bar left unfilled to mark what inference-time loss removes.
Fig. 2: Overview of the proposed method during training. Language and vision tokens are encoded by frozen pretrained encoders and integrated by a trainable adapter, while state tokens are projected and joined directly, so that the key/value memory of the action expert contains the tokens of all three modalities. Because the adapter attends across the language and vision tokens, each of its output tokens carries information from both, so the dependence on these two modalities is not measured on fully independent representations. The dependence pt drives two mechanisms: the tokens of mt−1† are zeroed in a subset of samples, and the entropy of pt enters the loss. In the step drawn here, vision is the target, so its token row is shaded and its tokens are dimmed. See Algorithm 1 for the order of operations.
Fig. 3: Top: modalities the policy receives from the bimanual manipulator CobotMagic: images from the head and the left- and right-hand cameras, taken at the moment when the slider of the fastener is grasped to close the zipper, and the motor currents. Bottom: six tasks used for real-world evaluation: bottle, close zipper, open zipper, towel folding, trash disposal, and unplug.
Fig. 4: Trajectory of the modality dependence pt,m during training, for each condition. The three modalities are overlaid in each panel; curves are exponential moving averages. The dashed horizontal line marks the uniform level 1/M≃0.33 , and the shaded band marks the interval in which the dropout probability ρt is nonzero (Baseline applies no dropout).
Fig. 5: Action delta at chunk arrival, in the normal setting and under vision loss, for each of the 30 trials per condition. Open and filled markers indicate the same trial, connected by a line, and horizontal bars are condition means. For both conditions, the action delta is larger when the chunk was generated under vision loss, with a smaller increase for TMD: the mean within-trial ratio of the two values, marked by × , is 3.12 for Baseline versus 2.29 for TMD. Action delta is the sum over the two arms of the L2 norm of the change in the commanded action vector between adjacent steps; the action vector comprises the six joint angles and the gripper command and is reported in arbitrary units.
Task
Baseline
TMD
RMD
bottle
2/5
5/5
0/5
close zipper
3/5
0/5
0/5
open zipper
1/5
2/5
0/5
towel folding
4/5
3/5
1/5
trash disposal
1/5
2/5
0/5
unplug
3/5
3/5
3/5
TABLE I: Number of successes / number of trials on the real robot under the normal setting ( N=5 per condition).
Task
Baseline ( Δ )
TMD ( Δ )
bottle
1/5 ( −1 )
5/5 ( 0 )
close zipper
0/5 ( −3 )
0/5 ( 0 )
open zipper
0/5 ( −1 )
2/5 ( 0 )
towel folding
1/5 ( −3 )
2/5 ( −1 )
trash disposal
0/5 ( −1 )
1/5 ( −1 )
unplug
0/5 ( −3 )
5/5 ( +2 )
TABLE II: Number of successes / number of trials on the real robot under vision loss ( N=5 per condition). Δ is the change from the normal setting.
Robotic systems perceive the world through multiple input modalities -- including visual camera streams and natural language instructions -- and must select appropriate actions based on these signals. However, assuming the permanent availability of all input devices is unrealistic, as sensors may fail, become occluded, or drop out entirely during deployment. Robust handling of such missing-modality scenarios is therefore essential for real-world robot operation. This paper introduces RL4IL, a reinforcement learning guided method for imitation learning that selects the most suitable action for a given observation by identifying the most relevant expert demonstrations from a training library. A reinforcement learning policy, trained via Proximal Policy Optimisation over Breadth-First Search candidate sets, ranks candidate demonstrations and a soft cross-attention fusion head aggregates their action signals to produce the final prediction. When a modality is missing at inference time, a dedicated per-modality RL retrieval policy identifies donor demonstrations from the training library, and a soft imputation head reconstructs the missing embedding via cross-attention over the top-ranked donors -- without requiring any retraining of the system. Experiments on three LIBERO benchmark suites demonstrate that RL4IL substantially outperforms state-of-the-art imitation learning methods under sensor dropout conditions, while requiring no policy network training. The code can be found at https://github.com/h-ismkhan/Reinforcement-Learning-via-kNN-for-Robotic-Learning-with-Missing-Camera
End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a tracking module. Experiments in simulation and the real world demonstrate that GuidedAttention consistently improves robot manipulation performance, particularly under positional and appearance out-of-distribution (OOD) conditions. https://mmurooka.github.io/guided-attention-project-page
Masaki Murooka, Ryoichi Nakajo, Keisuke Shirai +4
CNRS-AIST JRL (Joint Robotics Laboratory), IRL, 1-1-1 Umezono, Tsukuba, Ibaraki 305-8560, Japan · Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), 2-3-26 Aomi, Koto-ku, Tokyo 135-0064, Japan
Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of uninformative sensors and single-modality insufficiency when only one informative sensor remains intact. We propose Evidence-Gated Regularization (EGR), a modality-agnostic training objective that introduces zero inference-time overhead. EGR derives a per-frame and per-sensor task-relevance signal to gate two state-conditional consistency objectives: invariance on low-evidence sensors, and single-sensor sufficiency on high-evidence ones. We introduce a benchmark based on BEHAVIOR-1K, comprising a fast inference-only diagnostic suite and 47 rollout-based skills targeting modality entanglement. We validate EGR on this benchmark and on two real-robot setups with fundamentally different embodiments: a bi-manual setup with two Kinova arms and three RGB cameras, and a single-arm MELFA ASSISTA setup combining vision and GelSight tactile sensors. EGR improves simulation success rates (SR) from 12.5% to 16.4% under full modalities (+31%), from 9.4% to 16.5% under uninformative-sensor corruption (+75%), and from 2.8% to 6.1% under single-sensor fallback (+120%). Under physical-object distractors, EGR boosts SR from 30% to 85% on the bi-manual setup (+183%) and from 55% to 70% on the tactile setup (+27%).
Yue Yang, Diego Romeres, Chiori Hori +3
Department of Computer Science, University of North Carolina at Chapel Hill · 2Mitsubishi Electric Research Laboratories (MERL)