Organizations: Fujitsu Limited, Kanagawa 211-8588, Japan. · Faculty of Science and Engineering, Waseda University, Tokyo 169-8050, Japan. · National Institute of Advanced Industrial Science and Technology, Tokyo 100-8921, Japan.
Imitation learning policies that integrate multiple sensory modalities are prone to overreliance on a dominant modality, such as vision, during training, which can disrupt policy execution when that modality is lost at inference time. In this paper, we introduce Targeted Modality Dropout (TMD), in which the dependence on each modality is estimated using attention and the most dominant modality is selectively dropped. This is combined with entropy regularization over the dependence distribution. Through real-robot evaluation using a bimanual manipulator, we show that under vision loss the success rate of the baseline policy drops substantially, whereas TMD sustains task execution. In contrast, a conventional dropout that selects the dropped modality at random, without the entropy regularization, fails on many tasks even without vision loss.
Figures & tables
Fig. 1: Overview of the problem and of the proposed method. Three modalities feed one policy. (a) Under ordinary training the dependence, estimated from attention, concentrates on vision. (b) TMD drops whichever modality is dominant during training, and the dependence balances. The bars are the estimated dependence at the end of training, with the vision bar left unfilled to mark what inference-time loss removes.
Fig. 2: Overview of the proposed method during training. Language and vision tokens are encoded by frozen pretrained encoders and integrated by a trainable adapter, while state tokens are projected and joined directly, so that the key/value memory of the action expert contains the tokens of all three modalities. Because the adapter attends across the language and vision tokens, each of its output tokens carries information from both, so the dependence on these two modalities is not measured on fully independent representations. The dependence pt drives two mechanisms: the tokens of mt−1† are zeroed in a subset of samples, and the entropy of pt enters the loss. In the step drawn here, vision is the target, so its token row is shaded and its tokens are dimmed. See Algorithm 1 for the order of operations.
Fig. 3: Top: modalities the policy receives from the bimanual manipulator CobotMagic: images from the head and the left- and right-hand cameras, taken at the moment when the slider of the fastener is grasped to close the zipper, and the motor currents. Bottom: six tasks used for real-world evaluation: bottle, close zipper, open zipper, towel folding, trash disposal, and unplug.
Fig. 4: Trajectory of the modality dependence pt,m during training, for each condition. The three modalities are overlaid in each panel; curves are exponential moving averages. The dashed horizontal line marks the uniform level 1/M≃0.33 , and the shaded band marks the interval in which the dropout probability ρt is nonzero (Baseline applies no dropout).
Fig. 5: Action delta at chunk arrival, in the normal setting and under vision loss, for each of the 30 trials per condition. Open and filled markers indicate the same trial, connected by a line, and horizontal bars are condition means. For both conditions, the action delta is larger when the chunk was generated under vision loss, with a smaller increase for TMD: the mean within-trial ratio of the two values, marked by × , is 3.12 for Baseline versus 2.29 for TMD. Action delta is the sum over the two arms of the L2 norm of the change in the commanded action vector between adjacent steps; the action vector comprises the six joint angles and the gripper command and is reported in arbitrary units.
Task
Baseline
TMD
RMD
bottle
2/5
5/5
0/5
close zipper
3/5
0/5
0/5
open zipper
1/5
2/5
0/5
towel folding
4/5
3/5
1/5
trash disposal
1/5
2/5
0/5
unplug
3/5
3/5
3/5
TABLE I: Number of successes / number of trials on the real robot under the normal setting ( N=5 per condition).
Task
Baseline ( Δ )
TMD ( Δ )
bottle
1/5 ( −1 )
5/5 ( 0 )
close zipper
0/5 ( −3 )
0/5 ( 0 )
open zipper
0/5 ( −1 )
2/5 ( 0 )
towel folding
1/5 ( −3 )
2/5 ( −1 )
trash disposal
0/5 ( −1 )
1/5 ( −1 )
unplug
0/5 ( −3 )
5/5 ( +2 )
TABLE II: Number of successes / number of trials on the real robot under vision loss ( N=5 per condition). Δ is the change from the normal setting.
CNRS-AIST JRL (Joint Robotics Laboratory), IRL, 1-1-1 Umezono, Tsukuba, Ibaraki 305-8560, Japan · Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), 2-3-26 Aomi, Koto-ku, Tokyo 135-0064, Japan