Current grasp diffusion models provide rich priors for generation, yet their object-centric approach can violate the kinematic and collision constraints imposed by the embodiment and the environment. Existing embodiment-aware methods primarily perform local corrections around generated grasps through gradient guidance or optimization, making it difficult to recover from fundamentally infeasible modes. We present Steer2Grasp, a training-free, embodiment-agnostic framework for inference-time grasp steering that adapts a frozen Cartesian grasp diffusion model using deployment-specific rewards. Through Feynman-Kac (FK) inspired particle reweighting and resampling, the method reallocates population mass from infeasible to high-reward grasp modes, enabling population-level mode transitions without modifying the pretrained diffusion model or requiring differentiable constraints. The framework enables a unified treatment for single and dual arm grasping through reachability and collision aware rewards, followed by gradient free gripper level local refinement. Across diverse objects, robot embodiments, and constrained environments, our method substantially improves feasible grasp generation while maintaining proximity to the underlying grasp prior.
Figures & tables
Fig. 1 : Object-centric grasp models may generate force closure valid grasps that are unreachable or colliding for a specific robot embodiment configuration and environment. Local gradient guidance addresses this by perturbing individual denoising trajectories toward feasibility, but such local updates can degrade grasp quality and struggle to move between distant grasp modes. Steer2Grasp instead reweights and resamples the particle population, reallocating probability mass toward feasible modes while leaving the pretrained diffusion model unchanged.
Fig. 2 : Overview of the proposed method: Given an full object point cloud P and a pretrained Cartesian space grasp diffusion model, the reverse diffusion produces grasp hypotheses Ht−1 which are mapped to their clean grasp estimates H^0(Ht) using the Tweedie estimate, which is evaluated against deployment specific feasibility rewards. The resulting reward r(H^0;P,R,E) combine embodiment and environment information through multi-start IK, reachability and signed geometric clearances. These rewards instantiate the Feynman–Kac twisting potential used to re-weight and resample ( Hb ) the particle population in turn removing particles ( Ha ). Thus embodiment and environment constraints act through population level reweighting and resampling rather than modifying individual denoising transitions, enabling mode level steering toward feasible regions while retaining the multi modal structure of the pretrained grasp prior.
Fig. 3 : Evaluation settings. (a) Large Scale Evaluation. Top-down schematic of the scene generation scheme (see legend); the local frame at the object’s center shows its position and extent (w×d×h) , while θ and (x, y) at the sampled Panda arm(s) base give its heading and location relative to the world origin. The wall labels (NORTH / SOUTH / EAST / WEST) indicate the four candidate wall positions, of which kept vs. not-kept walls are distinguished per scene. (b) Curated Scenes. Top-Row: Single Arm Scenes from L-R: S1. Corridor , S2. Room Corner , S3. Shelf Top and S4. Partition Board . Bottom-Row: Dual Arm Scenes from L-R: S1. Table S2. Wall S3. Box S4. Shelf
Method
Reach. ↑
Coll. ↓
CS ↑
SR hold ↑
SR hold +CS ↑
Time ↓
VRAM ↓
%
%
@10
@all
%
@10
@all
s ( + base)
GB ( + base)
Single-Arm
Plain sampling
72.51
76.24
40.90
15.60
18.44
24.00
8.93
0.34 (base)
0.53 (base)
Rejection filter†
89.50
71.41
47.14
18.81
22.31
27.02
10.78
+1.12
+0.00
Oversampling
97.84
56.88
60.34
41.06
29.77
34.84
22.80
+1.52
+0.08
Gradient guidance
76.41
31.52
69.64
57.71
21.51
32.01
18.52
+11.23
+0.01
Ours, O
98.26
28.41
72.99
70.42
38.84
37.84
38.13
+3.32
+0.00
TABLE I : Large-scale scene evaluation. We generate N grasps for all methods; metrics are averaged across all scenes. Reach. is the reachability rate (IK convergence; for dual-arm, both arms must converge). Coll. refers to the collision rate. CS@k is constraint satisfaction among the top- k ranked grasps per scene. SR hold denotes stability under all pull axes. SR hold +CS@k is the joint fraction that is stable, reachable, and collision-free. Time and VRAM are end-to-end generation latency (s) and peak GPU memory (GB), reported as the change relative to plain sampling of the same setting; the plain-sampling row gives the absolute baseline. Rows below the rule are ours. Green / Orange indicate best/second-best per column within each setting. † Rejection filter on average generates fewer than N grasps per scene. Our method is denoted by shorthand O and with local refinement added to it, shorthand is O+LR .
Arm
Method
Single-Arm (SR task )
Dual-Arm (SR task )
S1
S2
S3
S4
S1
S2
S3
S4
Franka
Plain sampling
10/50
3/50
4/50
22/50
7/50
7/50
0/50
3/50
Rejection filter
10/29
3/34
4/27
22/37
7/26
7/16
0/14
3/9
Oversampling
48/50
13/50
29/50
49/50
39/50
45/50
1/50
9/50
Gradient Guidance
24/50
14/50
37/50
34/50
30/50
39/50
18/50
29/50
Ours, O
50/50
33/50
46/50
50/50
48/50
48/50
35/50
50/50
TABLE II : Curated-scene evaluation. SR task (successful/total grasps) on curated scenes (refer Fig. 3 ) for single- and dual-arm grasps across arm types. SR task is the fraction of generated grasps that are constraint-satisfying and successfully transport the object in MuJoCo . Green / Orange indicate best/second-best per column within each arm. Our method is denoted by shorthand O and with local refinement added to it, shorthand is O+LR .
Fig. 4 : Ablations : a) Effect of Temperature τ on Effective Sample Size and Constraint Satisfaction. b) Effect of rewards on Reachability and Collision
Fig. 5 : Real-world grasping. Columns correspond to the four objects; rows from top to bottom correspond to Easy , Medium , and Hard scenarios.
Arm Type
Object
Easy
Medium
Hard
Total
Single
Bowl
3/3
3/3
3/3
9/9
Cylinder
3/3
3/3
2/3
8/9
Dual
Bucket
3/3
2/3
1/3
6/9
Tray
3/3
2/3
2/3
7/9
TABLE III : Real-world success . 9 trials per object across varying difficulty.
Diffusion models sample effectively from high-dimensional, multimodal distributions, but their outputs may violate deployment constraints. For task-space robot policies, generated grasps, waypoints, or trajectories can be distributionally valid yet infeasible, violating reachability, collision-avoidance, or closed-loop executability requirements. This embodiment gap limits zero-shot deployment across robots, even when the task-space behavior itself is transferable. We propose an inference-time optimization framework that couples the behavior generation to physical feasibility by formulating diffusion guidance as a constrained optimization problem. Our key insight is to replace the sampling perturbation in the backward process with an optimized correction, allowing hard constraints or soft penalties to be imposed during sampling without the need to retrain the diffusion model, while keeping samples close to the learned prior. We evaluate the method on dexterous grasp synthesis with reachability and collision-avoidance constraints, and dynamic manipulation with controller-level trackability constraints. Across settings and robot embodiments, optimization-guided denoising matches the feasibility of projection- and gradient-guidance baselines while better preserving grasp quality, and improving controller-level executability and task success, with task success improving by up to 20pp. on dexterous grasping and 23pp. on visuomotor manipulation over the best baseline.
We study cross-embodiment 6-DOF robot grasping. Unlike prior works, we require the model not only to generalize to novel objects / scenes but also to novel gripper morphologies and physical grasping processes. Our method extends diffusion model based generative 6-DOF grasping models to condition on the additional gripper's representation. We propose a swept-volume heuristic for encoding the gripper. We train our cross-embodiment model with procedural grippers and a large-scale dataset of 2 Billion grasps. In simulation experiments, our model has the best zero-shot generalization to novel real-world grippers and objects over baseline methods. Our model also serves as a good initialization for fine-tuning to adapt to novel grippers. In ablations, we demonstrate the efficiency of our sweep-volume gripper representation and our procedural gripper training dataset. Last, we show zero-shot generalization to real-world novel grippers for 6-DOF grasping, surpassing baselines in cross-embodiment generalization.
Language-driven dexterous grasp models, such as DextER, perform well when instructions specify where to grasp, but we find they fail systematically when an instruction also specifies where not to grasp (e.g., "grasp the handle but avoid the body"). Existing training corpora, DexGYSNet among them, contain virtually no avoidance instructions, and collecting examples for every possible constraint is impractical. Moreover, because every part mentioned during training denotes a contact target, models may interpret a forbidden part as another region to grasp rather than one to avoid. We therefore introduce an inference-time framework for negation-constrained dexterous grasping that requires no negation-specific training examples. Combining Sequential Monte Carlo with classifier-free guidance, our method guides sampling toward the instructed part while pruning candidates headed for the forbidden region, without any negation examples during training. A frozen 3D part-grounding model localizes the forbidden region from the language instruction. To evaluate this setting, we construct NegGrasp, a benchmark of paired positive/negative instructions with constraint-aware metrics that credit a grasp only if it both accomplishes the task and respects the stated constraint. On NegGrasp, our method reduces the violation rate of the strongest baseline from 57.9% to 17.2% while improving both constraint-aware and physical success.
Geonho Kim, SooGon Kim, Jongmin Lee
Department of Computer Science & Engineering, Chung-Ang University