We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yielding strong data efficiency gains and robustness to out-of-distribution object positions and camera viewpoints. The action heads of current robot policies are typically formulated as an MLP regression from a single global feature vector produced by a pre-trained vision encoder. This global formulation requires the policy network to discover, from demonstrations alone, the relationship between target robot actions and the image features they project onto. The consequence is that although modern image features are semantically descriptive, spatially robust, and even multiview-consistent, the policies built on them are brittle to subtle changes in camera viewpoint and object placement--and surprisingly data-inefficient. BIND closes this gap by supplying the action-feature relationship through camera geometry rather than learning: it discretizes a volume of candidate end effector positions, attaches each candidate to the pre-trained features at its projection in each camera view, and selects actions by scoring each candidate's position and image-bound feature combination. On a real robot, we study data efficiency and out-of-distribution robustness to unseen object positions and camera viewpoints, as well as general long-horizon task execution and dexterity. We find BIND to be highly data-efficient and robust: it achieves near-perfect success on tasks with as few as 5 demonstrations, and degrades gracefully under steep camera-viewpoint shifts and held-out object positions where coordinate-regression baselines completely fail.
Figures & tables
Fig. 1 : Traditional global regression vs. BIND projection-anchored selection. Standard visuomotor policies (left) pool image features into a global descriptor and regress robot actions—forcing the network to discover the robot-action ↔ pixel mapping from demonstrations. BIND (right) instead discretizes a volume of candidate end effector positions and binds each candidate to the image features at its projection in each camera view via known camera poses and intrinsics. Intuitively, this offers the policy the information of ‘choosing this candidate robot action would move the robot EEF to this image feature.’
Fig. 2 : Data Efficiency and Out-of-Distribution Experiments: BIND vs. global-regression (ACT) and Motion-Tracks baselines. Across data efficiency (top), OOD object positions (middle), and OOD camera viewpoints (bottom), BIND’s projection-anchored robot action selection outperforms the global-regression ACT and Motion Tracks baselines. All heads share the same DINOv3 backbone and training data—the only difference is the action head.
Fig. 3 : Long-horizon and dexterous real-robot tasks. Rollouts across four tasks (left to right): teapot preparation, towel folding, spill cleanup, and cup stacking. Rows compare BIND against the ACT global-regression baseline and a fully-finetuned Molmo VLA; per-task bars report success rates for each method.
Fig. 4 : BIND forward pass. (1) An image encoder (DINOv3) encodes the posed scene and wrist views into pixel-aligned feature maps F1,F2 . (2) Each candidate robot action EEF point x in the discretized EEF volume is projected to its image features from both views, forming the action-feature pair [F1(x)∣F2(x)∣x] . (3) An MLP decodes each action-position combination into per-timestep probability volumes. (4) The per-timestep argmax over each probability volume yields the maximum-probability 3D EEF trajectory.
Method
Dual B.
Div. B.
Shoe
Cup
Apple
Mean
ManiFlow-3D ∗
54.0
72.3
68.3
72.7
42.0
61.9
ManiFlow-2D ∗
47.3
37.0
45.3
63.7
37.3
46.1
ACT
32.0
5.0
29.0
63.0
42.0
34.2
Diff. Policy
30.0
11.0
18.0
36.0
40.0
27.0
BIND (ours)
97.0
66.2
63.0
77.3
85.4
77.8
TABLE I : RoboTwin success rate (%) across five tasks. ∗ ManiFlow rows are quoted from their paper [ 25 ] ; our evaluation modifies the simulator viewpoint, so those rows are not strictly matched to ours. ManiFlow-3D consumes point clouds; ManiFlow-2D, ACT, and Diffusion Policy use RGB and proprioception.
BIND (ours)
ACT
Motion Tracks
a) Data efficiency: score vs. training-set size
3 demos
44
14
4
5
100
24
4
20
98
29
16
30
95
21
3
40
100
31
18
TABLE II : Real-robot cup pick-and-place.
Model
Spatial
Object
Avg
TraceVLA
84.6
85.2
84.9
OpenVLA
84.7
88.4
86.6
SpatialVLA
88.2
89.9
89.1
CoT-VLA
87.5
91.6
89.6
ThinkAct
88.3
91.4
89.9
MolmoAct-7B-D
87.0
95.4
91.2
TABLE III : LIBERO success rate (%) on the Spatial and Object suites. Rows sorted by the mean of the two suites.
In-dist.
OOD obj.
OOD view
BIND (ours)
94 / 94
96 / 96
49 / 44
Diffusion-x0
82 / 80
0 / 0
12 / 10
MolmoAct2
81 / 80
0 / 0
12 / 11
ACT
45 / 46
0 / 0
12 / 7
DP3
23 / 6
0 / 0
8 / 3
Motion Tracks
12 / 2
0 / 0
4 / 2
TABLE IV : Controlled simulation (MuJoCo cube pick-and-place). Cells report mean progress (0–100, partial credit for touching, picking, placing, and centering the cube) / binary success %; n=50 episodes for in-distribution and OOD object position, n=196 held-out cameras for OOD viewpoint (single training camera).
BIND
ACT
M. Tracks
MolmoAct2
Teapot (40 demos)
93
2
0
0
Cup stacking (40)
81
8
6
0
Spill cleanup (30)
100
0
3
0
Fold towel (35)
97
63
12
47
TABLE V : Long-horizon dexterous tasks (in-distribution). Mean progress score (0–100), including the MolmoAct2 baseline (fully-finetuned per task).
Backbone
Progress
Succ. %
DINOv3
98
82
DAv3
90
88
DynaFlip
84
80
PaliGemma
88
78
π0
78
60
TABLE VI : Backbone ablation (in-distribution pick-and-place). Same BIND head and training data with the pretrained vision encoder swapped. Progress is the mean 0–100 score; success is the binary pick-and-place rate.
Fig. 5 : Controlled simulation experiments (MuJoCo cube pick-and-place). We study data efficiency and OOD object position and OOD viewpoints with precise distribution definitions on a controlled MuJoCo cube pick-and-place setup. We compare BIND against five baselines (Diffusion Policy, MolmoAct2, ACT, Motion Tracks, and DP3). BIND is the only method that retains meaningful progress under both distribution shifts.