cs.ROSep 22, 2026

RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation

Authors: Han Qi, Heng Yang

Organizations: Harvard School of Engineering and Applied Sciences Harvard University

Abstract

Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.

Figures & tables

Explore similar work

CardsList
  1. Foundation and Small Models Coordination for Visuomotor Policy Learning

    Mar 9, 2026Haoran Ding, Liang Ma, Yaxun Yang +7Visuomotor PolicyRobotic Perception

  2. Action-Effect Memory Pretraining for Robot Manipulation

    Jun 10, 2026Yijing Zhou, Qiwei Liang, Sitong Zhuang +5Robotic ManipulationDual-Branch Temporal Modeling