We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yielding strong data efficiency gains and robustness to out-of-distribution object positions and camera viewpoints. The action heads of current robot policies are typically formulated as an MLP regression from a single global feature vector produced by a pre-trained vision encoder. This global formulation requires the policy network to discover, from demonstrations alone, the relationship between target robot actions and the image features they project onto. The consequence is that although modern image features are semantically descriptive, spatially robust, and even multiview-consistent, the policies built on them are brittle to subtle changes in camera viewpoint and object placement--and surprisingly data-inefficient. BIND closes this gap by supplying the action-feature relationship through camera geometry rather than learning: it discretizes a volume of candidate end effector positions, attaches each candidate to the pre-trained features at its projection in each camera view, and selects actions by scoring each candidate's position and image-bound feature combination. On a real robot, we study data efficiency and out-of-distribution robustness to unseen object positions and camera viewpoints, as well as general long-horizon task execution and dexterity. We find BIND to be highly data-efficient and robust: it achieves near-perfect success on tasks with as few as 5 demonstrations, and degrades gracefully under steep camera-viewpoint shifts and held-out object positions where coordinate-regression baselines completely fail.
Figures & tables
Fig. 1 : Traditional global regression vs. BIND projection-anchored selection. Standard visuomotor policies (left) pool image features into a global descriptor and regress robot actions—forcing the network to discover the robot-action ↔ pixel mapping from demonstrations. BIND (right) instead discretizes a volume of candidate end effector positions and binds each candidate to the image features at its projection in each camera view via known camera poses and intrinsics. Intuitively, this offers the policy the information of ‘choosing this candidate robot action would move the robot EEF to this image feature.’
Fig. 2 : Data Efficiency and Out-of-Distribution Experiments: BIND vs. global-regression (ACT) and Motion-Tracks baselines. Across data efficiency (top), OOD object positions (middle), and OOD camera viewpoints (bottom), BIND’s projection-anchored robot action selection outperforms the global-regression ACT and Motion Tracks baselines. All heads share the same DINOv3 backbone and training data—the only difference is the action head.
Fig. 3 : Long-horizon and dexterous real-robot tasks. Rollouts across four tasks (left to right): teapot preparation, towel folding, spill cleanup, and cup stacking. Rows compare BIND against the ACT global-regression baseline and a fully-finetuned Molmo VLA; per-task bars report success rates for each method.
Fig. 4 : BIND forward pass. (1) An image encoder (DINOv3) encodes the posed scene and wrist views into pixel-aligned feature maps F1,F2 . (2) Each candidate robot action EEF point x in the discretized EEF volume is projected to its image features from both views, forming the action-feature pair [F1(x)∣F2(x)∣x] . (3) An MLP decodes each action-position combination into per-timestep probability volumes. (4) The per-timestep argmax over each probability volume yields the maximum-probability 3D EEF trajectory.
Method
Dual B.
Div. B.
Shoe
Cup
Apple
Mean
ManiFlow-3D ∗
54.0
72.3
68.3
72.7
42.0
61.9
ManiFlow-2D ∗
47.3
37.0
45.3
63.7
37.3
46.1
ACT
32.0
5.0
29.0
63.0
42.0
34.2
Diff. Policy
30.0
11.0
18.0
36.0
40.0
27.0
BIND (ours)
97.0
66.2
63.0
77.3
85.4
77.8
TABLE I : RoboTwin success rate (%) across five tasks. ∗ ManiFlow rows are quoted from their paper [ 25 ] ; our evaluation modifies the simulator viewpoint, so those rows are not strictly matched to ours. ManiFlow-3D consumes point clouds; ManiFlow-2D, ACT, and Diffusion Policy use RGB and proprioception.
BIND (ours)
ACT
Motion Tracks
a) Data efficiency: score vs. training-set size
3 demos
44
14
4
5
100
24
4
20
98
29
16
30
95
21
3
40
100
31
18
TABLE II : Real-robot cup pick-and-place.
Model
Spatial
Object
Avg
TraceVLA
84.6
85.2
84.9
OpenVLA
84.7
88.4
86.6
SpatialVLA
88.2
89.9
89.1
CoT-VLA
87.5
91.6
89.6
ThinkAct
88.3
91.4
89.9
MolmoAct-7B-D
87.0
95.4
91.2
TABLE III : LIBERO success rate (%) on the Spatial and Object suites. Rows sorted by the mean of the two suites.
In-dist.
OOD obj.
OOD view
BIND (ours)
94 / 94
96 / 96
49 / 44
Diffusion-x0
82 / 80
0 / 0
12 / 10
MolmoAct2
81 / 80
0 / 0
12 / 11
ACT
45 / 46
0 / 0
12 / 7
DP3
23 / 6
0 / 0
8 / 3
Motion Tracks
12 / 2
0 / 0
4 / 2
TABLE IV : Controlled simulation (MuJoCo cube pick-and-place). Cells report mean progress (0–100, partial credit for touching, picking, placing, and centering the cube) / binary success %; n=50 episodes for in-distribution and OOD object position, n=196 held-out cameras for OOD viewpoint (single training camera).
BIND
ACT
M. Tracks
MolmoAct2
Teapot (40 demos)
93
2
0
0
Cup stacking (40)
81
8
6
0
Spill cleanup (30)
100
0
3
0
Fold towel (35)
97
63
12
47
TABLE V : Long-horizon dexterous tasks (in-distribution). Mean progress score (0–100), including the MolmoAct2 baseline (fully-finetuned per task).
Backbone
Progress
Succ. %
DINOv3
98
82
DAv3
90
88
DynaFlip
84
80
PaliGemma
88
78
π0
78
60
TABLE VI : Backbone ablation (in-distribution pick-and-place). Same BIND head and training data with the pretrained vision encoder swapped. Progress is the mean 0–100 score; success is the binary pick-and-place rate.
Fig. 5 : Controlled simulation experiments (MuJoCo cube pick-and-place). We study data efficiency and OOD object position and OOD viewpoints with precise distribution definitions on a controlled MuJoCo cube pick-and-place setup. We compare BIND against five baselines (Diffusion Policy, MolmoAct2, ACT, Motion Tracks, and DP3). BIND is the only method that retains meaningful progress under both distribution shifts.
The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A good choice of action representation and loss function can help to address these concerns, but there are often trade offs. We propose Action Map Policy (AMP), which casts 3D closed-loop manipulation policy learning as a classification problem in image space. While classification has been an effective formulation in generative language models, applying it to robot action learning is difficult because naively discretizing high-dimensional continuous actions explodes the token vocabulary. Our key idea is to project 3D actions onto the camera image planes and treat each pixel location as a discrete class, thus controlling dimensionality while retaining multi-modality. This method supports millimeter-level precision for high-dimensional actions without requiring a prohibitively large vocabulary, while preserving fine-grained pixel-wise visual signals. Furthermore, it can predict the entire action chunk in a single forward pass, avoiding complex noise scheduling and iterative denoising while achieving substantially faster inference than diffusion policies. Experiments on various manipulation tasks show that AMP outperforms strong baselines, achieving higher success rates, faster inference, and enhanced spatial reasoning.
Haojie Huang, Zhang Ye, Linfeng Zhao +7
Northeastern University · Stanford University · Brown University
Vision-language-action (VLA) models require 3D spatial reasoning, yet RGB observations encode robot-object geometry only implicitly. Lifting depth with camera intrinsics makes this geometry explicit as dense, image-aligned pointmaps, but their camera-frame coordinates depend on camera placement. We propose SeeR-VLA, which transforms pointmaps into a robot-centric frame with an end-effector origin and robot-base-aligned axes. An encoder initialized from pretrained RGB weights extracts pointmap features, which are added to corresponding RGB tokens without increasing the token count. Across 24 RoboCasa tasks and four real-world tasks, SeeR-VLA improves average success over RGB-only π0.5 by 6.4 and 32.5 percentage points, respectively. It also exceeds the strongest evaluated 3D-augmented baseline, PointVLA, by 3.5 and 23.7 percentage points, respectively. Beyond these gains, our ablations clarify how coordinate choices affect VLA performance, showing that end-effector centering is most effective with robot-base-aligned axes. The benefits grow as training viewpoints diversify, highlighting the importance of using robot-frame pointmaps when learning from diverse camera configurations.
Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliably execute correct low-level actions under distribution shifts. For example, even in a simple pickup task with identical scene layouts, camera viewpoints, and illumination, performance can degrade substantially when the object is placed at unseen locations. We argue that this gap arises from insufficient action understanding, namely the inability to interpret the robot's base-frame action coordinate system in image space. To address this issue, we introduce AxisGuide, a lightweight guidance method that bridges semantic scene understanding and action-coordinate interpretation. Using camera parameters and end-effector poses, AxisGuide renders the robot base-frame axes in each camera view and augments RGB observations with a small set of cue channels that explicitly visualize the meaning of the +x, +y, and +z motions in image space. Extensive evaluations in both the LIBERO simulation and real-world environments demonstrate that AxisGuide yields substantial performance gains and improved generalization, highlighting the effectiveness of explicit action-coordinate cues for learning reliable and transferable generalist visuomotor policies.
Jiyun Jang, Yujin Sung, Woosung Joung +5
1Korea University · University of Michigan · 3KT R&D Center +1