Current grasp diffusion models provide rich priors for generation, yet their object-centric approach can violate the kinematic and collision constraints imposed by the embodiment and the environment. Existing embodiment-aware methods primarily perform local corrections around generated grasps through gradient guidance or optimization, making it difficult to recover from fundamentally infeasible modes. We present Steer2Grasp, a training-free, embodiment-agnostic framework for inference-time grasp steering that adapts a frozen Cartesian grasp diffusion model using deployment-specific rewards. Through Feynman-Kac (FK) inspired particle reweighting and resampling, the method reallocates population mass from infeasible to high-reward grasp modes, enabling population-level mode transitions without modifying the pretrained diffusion model or requiring differentiable constraints. The framework enables a unified treatment for single and dual arm grasping through reachability and collision aware rewards, followed by gradient free gripper level local refinement. Across diverse objects, robot embodiments, and constrained environments, our method substantially improves feasible grasp generation while maintaining proximity to the underlying grasp prior.
Figures & tables
Fig. 1 : Object-centric grasp models may generate force closure valid grasps that are unreachable or colliding for a specific robot embodiment configuration and environment. Local gradient guidance addresses this by perturbing individual denoising trajectories toward feasibility, but such local updates can degrade grasp quality and struggle to move between distant grasp modes. Steer2Grasp instead reweights and resamples the particle population, reallocating probability mass toward feasible modes while leaving the pretrained diffusion model unchanged.
Fig. 2 : Overview of the proposed method: Given an full object point cloud P and a pretrained Cartesian space grasp diffusion model, the reverse diffusion produces grasp hypotheses Ht−1 which are mapped to their clean grasp estimates H^0(Ht) using the Tweedie estimate, which is evaluated against deployment specific feasibility rewards. The resulting reward r(H^0;P,R,E) combine embodiment and environment information through multi-start IK, reachability and signed geometric clearances. These rewards instantiate the Feynman–Kac twisting potential used to re-weight and resample ( Hb ) the particle population in turn removing particles ( Ha ). Thus embodiment and environment constraints act through population level reweighting and resampling rather than modifying individual denoising transitions, enabling mode level steering toward feasible regions while retaining the multi modal structure of the pretrained grasp prior.
Fig. 3 : Evaluation settings. (a) Large Scale Evaluation. Top-down schematic of the scene generation scheme (see legend); the local frame at the object’s center shows its position and extent (w×d×h) , while θ and (x, y) at the sampled Panda arm(s) base give its heading and location relative to the world origin. The wall labels (NORTH / SOUTH / EAST / WEST) indicate the four candidate wall positions, of which kept vs. not-kept walls are distinguished per scene. (b) Curated Scenes. Top-Row: Single Arm Scenes from L-R: S1. Corridor , S2. Room Corner , S3. Shelf Top and S4. Partition Board . Bottom-Row: Dual Arm Scenes from L-R: S1. Table S2. Wall S3. Box S4. Shelf
Method
Reach. ↑
Coll. ↓
CS ↑
SR hold ↑
SR hold +CS ↑
Time ↓
VRAM ↓
%
%
@10
@all
%
@10
@all
s ( + base)
GB ( + base)
Single-Arm
Plain sampling
72.51
76.24
40.90
15.60
18.44
24.00
8.93
0.34 (base)
0.53 (base)
Rejection filter†
89.50
71.41
47.14
18.81
22.31
27.02
10.78
+1.12
+0.00
Oversampling
97.84
56.88
60.34
41.06
29.77
34.84
22.80
+1.52
+0.08
Gradient guidance
76.41
31.52
69.64
57.71
21.51
32.01
18.52
+11.23
+0.01
Ours, O
98.26
28.41
72.99
70.42
38.84
37.84
38.13
+3.32
+0.00
TABLE I : Large-scale scene evaluation. We generate N grasps for all methods; metrics are averaged across all scenes. Reach. is the reachability rate (IK convergence; for dual-arm, both arms must converge). Coll. refers to the collision rate. CS@k is constraint satisfaction among the top- k ranked grasps per scene. SR hold denotes stability under all pull axes. SR hold +CS@k is the joint fraction that is stable, reachable, and collision-free. Time and VRAM are end-to-end generation latency (s) and peak GPU memory (GB), reported as the change relative to plain sampling of the same setting; the plain-sampling row gives the absolute baseline. Rows below the rule are ours. Green / Orange indicate best/second-best per column within each setting. † Rejection filter on average generates fewer than N grasps per scene. Our method is denoted by shorthand O and with local refinement added to it, shorthand is O+LR .
Arm
Method
Single-Arm (SR task )
Dual-Arm (SR task )
S1
S2
S3
S4
S1
S2
S3
S4
Franka
Plain sampling
10/50
3/50
4/50
22/50
7/50
7/50
0/50
3/50
Rejection filter
10/29
3/34
4/27
22/37
7/26
7/16
0/14
3/9
Oversampling
48/50
13/50
29/50
49/50
39/50
45/50
1/50
9/50
Gradient Guidance
24/50
14/50
37/50
34/50
30/50
39/50
18/50
29/50
Ours, O
50/50
33/50
46/50
50/50
48/50
48/50
35/50
50/50
TABLE II : Curated-scene evaluation. SR task (successful/total grasps) on curated scenes (refer Fig. 3 ) for single- and dual-arm grasps across arm types. SR task is the fraction of generated grasps that are constraint-satisfying and successfully transport the object in MuJoCo . Green / Orange indicate best/second-best per column within each arm. Our method is denoted by shorthand O and with local refinement added to it, shorthand is O+LR .
Fig. 4 : Ablations : a) Effect of Temperature τ on Effective Sample Size and Constraint Satisfaction. b) Effect of rewards on Reachability and Collision
Fig. 5 : Real-world grasping. Columns correspond to the four objects; rows from top to bottom correspond to Easy , Medium , and Hard scenarios.
Arm Type
Object
Easy
Medium
Hard
Total
Single
Bowl
3/3
3/3
3/3
9/9
Cylinder
3/3
3/3
2/3
8/9
Dual
Bucket
3/3
2/3
1/3
6/9
Tray
3/3
2/3
2/3
7/9
TABLE III : Real-world success . 9 trials per object across varying difficulty.