We propose NUDGE (Nudge Update via Differentiable GEometry), a training-free obstacle-avoidance procedure that can be incorporated in any robot policy based on diffusion or flow matching, including diffusion policies and vision-language-action models. Our work injects gradients from a signed distance field, a function returning each point's distance to the nearest obstacle, into the policy at inference time to steer it away from obstacles. It supports any common action parameterization, from absolute or relative joint poses to end-effector poses, through a differentiable joint-trajectory decoder. Experiments show that NUDGE preserves the policy's task distribution and runs reactively in real time.
Figures & tables
Fig. 1: NUDGE provides zero-shot collision avoidance for generative robot policies at inference time. (a) From a live point cloud, NUDGE computes a signed distance field in real time to react to scene changes, injects its gradient into the policy’s denoising loop, and steers the predicted action chunk away from obstacles. (b) We fine-tuned a π0 policy for a table-cleaning task, where the policy learns to model the contact and task distribution from demonstrations. Top: the unguided baseline collides with the pool noodle. Middle & Bottom: NUDGE avoids the obstacle at the extra cost of only 2.5 ms per action chunk.
Action representation
Notation
Decoder qt+i=Di(A;qt)
Joint pose
ai=qitgt
qt+i=ai
Joint delta
ai=Δqi
qt+i=qt+∑j=1iaj
EEF pose
ai=Tiee
qt+i=IK(ai;qt+i−1)
EEF delta
ai=ΔTiee
qt+i=IK(Ttee⋅∏j=1iaj;qt+i−1)
TABLE I: Action-space decoders
Goal
Collision
Manifold
Method
Success ↑
Reach ↑
Coll. Ep. ↓
Coll. Rate ↓
On-Sphere ↑
Sphere Err. (cm) ↓
Vanilla
37.0%
82.0%
54.0%
16.4%
84.3%
2.78
Post-Projection
25.0%
29.0%
14.5%
4.7%
56.4%
4.84
OmniGuide [ 37 ]
39.0%
81.0%
52.0%
15.98%
84.9%
2.71
RAIL [ 4 ]
37.5%
37.5%
0.0%
0.00%
82.9%
2.73
NUDGE, ρk=0.01
42.0%
73.0%
47.0%
13.8%
80.5%
3.00
TABLE II: Sphere-manifold results
Fig. 2: Collision-query geometry: NUDGE’s whole-body spheres versus OmniGuide’s EEF point and four wrist probe.
Fig. 3: Left. The robot trajectory from start to goal with end-effort constrained on the sphere manifold. Right. Failure modes on the sphere constraint task: collides with the obstacle and pushes the end-effector off the learned task manifold.
Fig. 4: One- and three-obstacle LIBERO-Object scenes with matched initial robot/object poses and camera view. We note that the additional obstacles may make a collision-free grasp infeasible.
Unguided
OmniGuide
RAIL
NUDGE
Task
SR ↑
CR ↓
SR ↑
CR ↓
SR ↑
CR ↓
SR ↑
CR ↓
One obstacle
Can
50
5.70
60
15.19
10
0.00
100
0.00
Cheese
40
6.56
40
17.25
0
0.00
60
9.71
Dressing
20
4.55
40
0.09
0
0.00
40
0.00
BBQ
0
4.72
10
2.22
0
0.00
10
0.03
TABLE III: LIBERO-Object results
Fig. 5: Top: Denoising visualization for a single action chunk. The policy predicts a 50 -step chunk of joint deltas. The actions are mapped to end-effector pose and rendered in a color spectrum across denoising steps, from random noise at t=10 down to the final denoised chunk at t=0 . The environment is shown as a voxel map. Bottom, left: Real-world setup; the obstacle is a pool noodle. Bottom, middle and right: Final action chunk at t=0 for the unguided baseline and NUDGE, respectively. The baseline is unaware of the obstacle and collides with it. NUDGE pushes the chunk away from the obstacle.
Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-world scenes with unstructured obstacles remains a key generalization challenge. We present ObstaDiff, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder. ObstaDiff extracts a structured target-obstacle-background representation, enabling the downstream alignment policy to generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles. We evaluate ObstaDiff on 61 real-robot greenhouse trials per method (366 executions in total). ObstaDiff achieves 75.41% average task success and 8.20% average obstacle collision rate, outperforming representative imitation-learning baselines and improving generalization in cluttered agricultural scenes.
Jiawen Wang, Kevin Yao, Khalid Jawed
Department of Mechanical and Aerospace Engineering, University of California, Los Angeles · Department of Computer Science, University of California, Los Angeles
Diffusion models sample effectively from high-dimensional, multimodal distributions, but their outputs may violate deployment constraints. For task-space robot policies, generated grasps, waypoints, or trajectories can be distributionally valid yet infeasible, violating reachability, collision-avoidance, or closed-loop executability requirements. This embodiment gap limits zero-shot deployment across robots, even when the task-space behavior itself is transferable. We propose an inference-time optimization framework that couples the behavior generation to physical feasibility by formulating diffusion guidance as a constrained optimization problem. Our key insight is to replace the sampling perturbation in the backward process with an optimized correction, allowing hard constraints or soft penalties to be imposed during sampling without the need to retrain the diffusion model, while keeping samples close to the learned prior. We evaluate the method on dexterous grasp synthesis with reachability and collision-avoidance constraints, and dynamic manipulation with controller-level trackability constraints. Across settings and robot embodiments, optimization-guided denoising matches the feasibility of projection- and gradient-guidance baselines while better preserving grasp quality, and improving controller-level executability and task success, with task success improving by up to 20pp. on dexterous grasping and 23pp. on visuomotor manipulation over the best baseline.
Diffusion policies generate multimodal robot action sequences from demonstrations, but steering them toward deployment-time constraints typically relies on differentiable guidance costs. This excludes many practical safety constraints, such as binary collision checks, joint limits, and black-box rollout costs that are nondifferentiable. We propose Gradient-free Robot Action generation via Combined diffusion-MPPI posterior mean Estimation (GRACE), which guides a pretrained diffusion policy with Model Predictive Path Integral (MPPI) control using only forward cost evaluations. Building on the common score-ascent structure of diffusion and MPPI, GRACE constructs a cost-conditioned guidance posterior at each reverse step and estimates its mean with a single MPPI update centered at the diffusion reverse mean. For differentiable costs, GRACE recovers conventional gradient guidance under a first-order, matched-covariance approximation. GRACE attains higher success rates than diffusion-based and sampling-based baselines in simulation. On a real 7-DoF manipulator, GRACE avoids a deployment-time obstacle that the unguided prior collides with in every trial. Code and experiment videos are available at https://anonymous.4open.science/w/grace-70BB/.
Leesai Park, Jiho HOng, Sanghyun Kim
School of Mechanical Engineering, Kyung Hee University, Yongin, Republic of Korea. · Advanced Institute of Convergence Technology (AICT), Suwon, Republic of Korea.