When humans learn new manipulation skills, they are able to generalize these skills to new contexts and environments. In particular, when learning, humans can easily separate task-relevant aspects of the environment (e.g., object location) and distractors (e.g., the table color). Ideally, robot policies should be able to learn similar, generalizable representations, but instead they often fail under small environment shifts such as changes in lighting, object instance, or initial configuration. In this paper, we propose a self-supervised method that learns a mask which, when multiplied by the observed image features, attempts to transform these features to retain only those which are relevant to the task. Our method --- which we call TransMASK --- can be combined with a variety of imitation learning frameworks (such as diffusion policies) without any additional labels or alterations to the loss function. By introducing a learned mask to the network during training, we aim to induce competitive pressure among the image features during training to force the policy to only attend to features which are consistently task-relevant. We find empirically that our masks are interpretable and can reject known spurious features such as the position of distractor objects or the background color. When compared to other representation learning methods for imitation learning, we find that TransMASK results in policies that are more robust to distribution shifts for irrelevant features, achieving at least 30 % improvement over the baselines when tested on out-of-distribution environments. See our project website: https://transmask.github.io/TransMASK/
Figures & tables
Fig. 1: Robot learning to pick up a green block from a cluttered environment and place it at the center of the table. (A) The expert demonstrations are collected over a wooden table. (B) This expert’s decisions depend only on features intrinsic to the task (e.g., green block, robot position, and target position). The robot’s observations in these demonstrations, however, record information about the entire scene, capturing the table texture, background color, and task-irrelevant objects. (C) We assume a disentangled state s is extracted from these high-dimensional visual observations. In this state, some elements correspond to the task structure — positions of the green block, target, and robot, while others correspond to scene-specific factors. In the figure, relevant elements are shown in green and scene-specific elements are shown in orange and purple. (D) Standard imitation learning policy attends to the entire state. When this policy is deployed for the identical task in a different scene (e.g., over a marble table), it may fail due to spurious dependencies learned during training. (E) To learn a robust policy, we must encode the state to a representation z which masks-out task-irrelevant features. A policy that is conditioned on z is more robust to distribution shift caused by changes in the scene that do not alter the task structure, for instance, when the task is still to pick and place a green block but over a marble table instead of a wooden table.
Fig. 2: Schematic diagram of TransMASK. The visual observations are encoded to a disentangled state vector s . We introduce a mask encoder that outputs a constant mask M of shape n×n , where n is the dimension of s . This mask is a sparse matrix in which the columns that correspond to task-irrelevant elements of s are close to zero. Therefore, when we compute the representation z by transforming the state with M , it only retains the elements critical for accurately predicting actions. The robot policy is conditioned on z rather than the entire state to achieve robustness.
Fig. 3: Overview of how the mask is learned. (Left) The mask encoder inputs a 1 -vector of size k ( 1k=[1,1,⋯,1]∈Rk ) and outputs a matrix M∈Rn×n . Since the expert’s actions when providing demonstrations are only influenced by the task-relevant features in the environment, the columns of the Jacobian of the expert policy corresponding to the irrelevant elements in the state will be near-zero magnitude as discussed in Section IV-B . (Right) When training the robot policy, as the loss converges the Jacobian of the robot’s learned policy changes the rows of M until the values start weighting the task-relevant elements more than the irrelevant ones.
Fig. 4: Results from our simulated experiments. We perform the simulated experiments in three tasks — Pick , Push , and Rotate . The methods are trained with two different policy heads: a fully connected policy (MLP) and a diffusion policy (DP). We evaluate the methods in two scenes, denoted as In-Distribution or ID (over wooden table, top row), and Out-of-Distribution or OOD (over marble table, bottom row). All methods are trained on demonstrations collected in the ID scene, except VINN and CLASS which are trained on demonstratinos from both ID and OOD scenes. We use the Compact Letter Display (CLD) algorithm [ 44 ] which assigns letters to each policy. The policies which do not share a letter have statistically significant difference in their results.
Fig. 5: Learned mask over the course of training. Our method learns a mask M∈Rn×n , and computes the latent state as z=Ms . Therefore, the magnitude of the i -th column of M corresponds to the ith element of s . Here, we show how the weights of the state change as the policy training progresses. We label which components of the column relate to which information in the scene. The components with a dark blue shade have values closer to 0 , indicating that the corresponding feature of the scene is removed from z . In contrast, the components with a red shade have values closer to 1 , indicating that the associated feature is retained in z .
Fig. 6: Results of our real-world experiments. We test the methods in three tasks — Pick , Stack , and Scoop . We evaluate the methods in two scenes, denoted as In-Distribution or ID (over wooden table), Out-of-Distribution or OOD (the table is covered with a white sheet). All the methods are trained on demonstrations collected over the wooden table, except VINN and CLASS which are trained on a mixed dataset with demonstrations collected from ID and OOD scenes. The policies which do not share a letter have statistically significant difference in the results.
Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effective states from high-dimensional image inputs and limited samples for sample-efficient reinforcement learning. To address this challenge, inspired by fields such as natural language processing and computer vision, we propose a self-supervised task based on mask prediction as an auxiliary task for reinforcement learning. This non-reconstruction method uses the sequence information collected by the agent from the environment and the context information in the sequence to predict the masked information, thereby strengthening the agent's understanding of the task and learning effective representations. Combined with transformers, we find that the model reconstructs the masked input sequence in the latent space. By feeding the compressed representations learned by this method into reinforcement learning models, we observe an improvement in the sample efficiency of reinforcement learning. Moreover, the model outperforms state-of-the-art sample-efficient reinforcement learning methods on multiple continuous and discrete control benchmarks.
Kai Zhao
School of Systems Science, Beijing Normal University · Beijing, China
Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.
Han Qi, Heng Yang
Harvard School of Engineering and Applied Sciences Harvard University
Visuomotor policies trained via behavior cloning are vulnerable to covariate shift, where small deviations from expert trajectories can compound into failure. Common strategies to mitigate this issue involve expanding the training distribution through human-in-the-loop corrections or synthetic data augmentation. However, these approaches are often labor-intensive, rely on strong task assumptions, or compromise the quality of imitation. We introduce Latent Policy Barrier, a framework for robust visuomotor policy learning. Inspired by Control Barrier Functions, LPB treats the latent embeddings of expert demonstrations as an implicit barrier separating safe, in-distribution states from unsafe, out-of-distribution (OOD) ones. Our approach decouples the role of precise expert imitation and OOD recovery into two separate modules: a base diffusion policy solely on expert data, and a dynamics model trained on both expert and suboptimal policy rollout data. At inference time, the dynamics model predicts future latent states and optimizes them to stay within the expert distribution. Both simulated and real-world experiments show that LPB improves both policy robustness and data efficiency, enabling reliable manipulation from limited expert data and without additional human correction or annotation.