RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation
Organizations: Harvard School of Engineering and Applied Sciences Harvard University
Abstract
Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.
Figures & tables
| Method | Place cube | Put can (clear) | Put can (obstacle) | Pull tool | Peg insertion | Long horizon | Two- arm | Avg. |
|---|---|---|---|---|---|---|---|---|
| Clean Spatial OOD | ||||||||
| 2D object-centric masking | .350 | .190 | .000 | .000 | .020 | .007 | .000 | .081 |
| 3D object-centric masking | .850 | .840 | .070 | .000 | .085 | .036 | .120 | .286 |
| DP–DINOv2 | .410 | .820 | .750 | .150 | .170 | .812 | .440 | .507 |
| RoboGround–DINOv2 | .690 | .820 | .770 | .000 | .215 | .789 | .630 | .559 |
| RoboMP-DINOv2 | .540 | .860 | .830 | .170 | .230 | .848 | .770 | .607 |
| Cosine | Relative | |||||
|---|---|---|---|---|---|---|
| Task | Shift | DP–DINOv2 | RoboMP-DINOv2 / -MCR | DP–DINOv2 | RoboMP-DINOv2 / -MCR | |
| Out-of-bin | Clutter | 1269 | ||||
| Appearance | 1269 | |||||
| Pull tool | Clutter | 2409 | ||||
| Appearance | 2409 | |||||
| Peg insertion | Clutter | 1514 | ||||
| Method | Place cube | Put can (clear) | Put can (obstacle) | Pull tool | Peg insertion | Long horizon | Two- arm | Avg. |
|---|---|---|---|---|---|---|---|---|
| 2D object-centric masking (masked-area rand.) | .790 | .040 | .010 | .000 | .000 | .003 | .000 | .120 |
| 3D object-centric masking (masked-area rand.) | .940 | .030 | .000 | .000 | .010 | .015 | .000 | .142 |
| DP–DINOv2 (full-image rand.) | .330 | .050 | .100 | .100 | .250 | .498 | .010 | .191 |
| RoboGround–DINOv2 (masked-area rand.) | .990 | .650 | .660 | .000 | .000 | .160 | .000 | .351 |
| RoboMP-DINOv2-MCR | .990 | .710 | .670 | .820 | .700 | .906 | .280 | .725 |
| Method | Two arm | Place cube | Pull tool |
|---|---|---|---|
| DP + full-image rand. | 0.00 | 0.02 | 0.00 |
| DP + masked rand. | 0.14 | 0.08 | 0.34 |
| Only mask + mask rand. | 0.00 | 0.04 | 0.00 |
| RoboMP-DINOv2-MCR | 0.26 | 0.24 | 0.38 |