Model-Based Geometry-Aware Generative Optimization for Constrained Locomotion Planning
Organizations: Carnegie Mellon University · University of Michigan, Ann Arbor
Abstract
Constrained Locomotion Planning (CLP) for quadrupeds and humanoids, where robots must satisfy collision avoidance, contact consistency, kinematic feasibility, and support constraints, is challenging under high-dimensional dynamics and highly non-convex environments. Recent Model-Based Diffusion (MBD) approaches recast trajectory optimization as posterior sampling over trajectories, using known dynamics and Monte Carlo rollouts to analytically estimate the denoising score function without demonstration learning. While constrained variants further incorporate feasibility into model-based score rollouts and show promising performance, they are still limited by (1) lacking a task-modulated active constraint geometry that shapes the score direction and reverse stochasticity, and (2) using deterministic DDPM-style reverse transport without adaptive scheduling across different generative transports. Therefore, we introduce Model-Based Geometry-Aware Generative Optimization (2GO) for constrained locomotion, which turns active constraint geometry into executable denoising operators through normal- induced metric shaping, tangent-space stochastic filtering, and CFS-based retraction. 2GO further decouples generative transport from reverse stochasticity through an adaptive diffusion and flow-like schedule. Experiments on constrained quadruped and humanoid locomotion demonstrate strong performance in discrete foothold selection and continuous posture planning, with higher success rates, fewer violations, and improved execution compatibility.
Figures & tables
| Env. | MPPI | MBD | MDOC | MD-COAS | 2GO | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| pSSR | eSSR | pSSR | eSSR | pSSR | eSSR | pSSR | eSSR | pSSR | eSSR | |||||||||||
| S1 | ||||||||||||||||||||
| S2 | ||||||||||||||||||||
| ZA | ||||||||||||||||||||
| ZB | ||||||||||||||||||||
| ZC | ||||||||||||||||||||
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | State and action summary | ||||
|---|---|---|---|---|---|
| Stepping stones | 16 | 12 | 16 | 1.0 | ; . |
| Humanoid corridor | 14 | 9 | 50 | 0.25 | ; . |
| Zone | Start goal | Constraint geometry |
|---|---|---|
| A | Centered thin wall, , . | |
| B | Connected U-wall with top segment , . | |
| C | Tapered squeeze over with a main gap and quarter-circle entry and exit. | |
| D | Two full-height planning obstacles centered at and , each with radius . |
| Component | Value or behavior |
|---|---|
| Quadruped | Interior clearance , stride , maximum swing , touchdown tolerance , simulation step ; leg order ; phase length 36 simulation steps. |
| Base assistance | Planar PD gains , height gains , roll/pitch gains , yaw gains ; direct base wrench scaled by the support-contact count. |
| Evaluation timing | Interface 50 Hz; simulation step s; warmup s; reference speed factor ; terminal hold 5 s. |
| Reference feedback | Planar gain , correction cap m/s; yaw feedback gain ; planar low-pass weight . |
| Adapter speed box | World- and world- m/s; local step s. |
| Command slew | No explicit bound on successive locomotion-command differences. |
| Quantity | Planning model | Reference adapter | Tracker / policy |
|---|---|---|---|
| Command magnitude | Humanoid body-frame velocity and posture-rate bounds; control-magnitude penalties. | World-frame position-increment bounds specify reference speeds; safety-box fallback can relax nominal bounds. | Body-frame planar velocity and separately generated yaw-rate commands are clipped at the policy input. |
| Command slew | Control-magnitude penalties do not constrain successive-command differences. | Retiming and low-pass filtering moderate command variation; the reference-speed box is not a slew bound. | Input saturation bounds magnitude; no explicit command-slew limit is imposed. |
| Posture | Bounds on height, torso/arm posture, and their command rates. | Planned posture enters body clearance; a pelvis-roll proxy adds a lateral-speed bound, not a measured-roll constraint. | Upper-body targets follow a separate mapping/PD path; the learned policy produces leg targets, followed by joint-limit clipping. |
| Clearance | Task geometry and planning/CFS margins are evaluated along the model rollout. | The corridor wall box and local lateral body-clearance half-space use ; finite correction can retain residuals. | Filtering and saturation follow adaptation without a final-reference clearance recheck; executed proxy clearance is evaluated afterward. |
| Quadruped footholds and support | Foot-placement geometry, support-offset bounds, and support-proximity cost. | Gait construction, stance anchors, and reachable touchdown projection; support-polygon membership is diagnostic. | Foot tracking and joint PD execute the gait with the common supplementary floating-base wrench. |
| pSSR | (%) | (s) | ||||
| Seed | CFS on | CFS off | CFS on | CFS off | CFS on | CFS off |
| 0 | 0.60 | 0.30 | 2.01 | 1.25 | 111.15 | 64.33 |
| 1 | 0.70 | 0.40 | 0.50 | 0.67 | 77.40 | 34.31 |
| 2 | 0.40 | 0.20 | 1.93 | 1.76 | 58.54 | 27.84 |
| 3 | 0.60 | 0.30 | 1.44 | 1.17 | 60.49 | 49.19 |
| 4 | 0.50 | 0.30 | 2.45 | 1.53 | 63.07 | 37.58 |
| Method | Seed | pSSR | Goal error (m) | Foot distance (m) | Fall | Success | ||
|---|---|---|---|---|---|---|---|---|
| MPPI | 0 | 1.00 | 2.232 | 3.538 | 0.0622 | 0.0521 | — | 0 |
| 1 | 0.00 | 2.236 | 2.750 | 0.0792 | 0.0526 | — | 0 | |
| 2 | 1.00 | 2.300 | 2.626 | 0.1364 | 0.0477 | — | 1 | |
| 3 | 0.00 | 1.932 | 2.652 | 0.0974 | 0.0615 | — | 0 | |
| 4 | 0.00 | 2.314 | 2.761 | 0.0837 | 0.0520 | — | 0 | |
| MBD | 0 | 0.30 | 2.736 | 17.594 | 0.0512 | 0.0454 | — | 1 |
| Method | Seed | pSSR | Goal error (m) | Foot distance (m) | Fall | Success | ||
|---|---|---|---|---|---|---|---|---|
| MPPI | 0 | 0.00 | 2.719 | 2.647 | 0.0713 | 0.0533 | — | 0 |
| 1 | 0.00 | 2.559 | 2.693 | 0.0823 | 0.0510 | — | 0 | |
| 2 | 1.00 | 2.445 | 2.231 | 0.0814 | 0.0550 | — | 0 | |
| 3 | 0.00 | 3.133 | 2.794 | 0.1479 | 0.0594 | — | 0 | |
| 4 | 0.00 | 2.873 | 2.263 | 0.1402 | 0.0583 | — | 0 | |
| MBD | 0 | 0.00 | 2.781 | 16.041 | 0.0639 | 0.0489 | — | 1 |
| Method | Seed | pSSR | Goal error (m) | Min. proxy SDF (m) | Fall | Success | ||
|---|---|---|---|---|---|---|---|---|
| MPPI | 0 | 0.00 | 0.667 | 1.966 | 0.1618 | 0.0090 | 0 | 1 |
| 1 | 0.00 | 0.750 | 1.163 | 0.1610 | 0.0074 | 0 | 1 | |
| 2 | 0.00 | 1.333 | 1.136 | 0.1909 | -0.0105 | 0 | 0 | |
| 3 | 0.00 | 1.250 | 1.476 | 0.4769 | -0.0395 | 0 | 0 | |
| 4 | 0.00 | 1.750 | 1.244 | 0.2070 | -0.0615 | 0 | 0 | |
| MBD | 0 | 0.00 | 1.067 | 12.037 | 0.1056 | -0.0525 | 0 | 0 |
| Method | Seed | pSSR | Goal error (m) | Min. proxy SDF (m) | Fall | Success | ||
|---|---|---|---|---|---|---|---|---|
| MPPI | 0 | 0.00 | 3.333 | 2.087 | 0.2071 | -0.0433 | 0 | 0 |
| 1 | 0.00 | 3.917 | 1.140 | 0.4910 | -0.1241 | 0 | 0 | |
| 2 | 0.00 | 3.917 | 1.177 | 0.4753 | -0.2029 | 0 | 0 | |
| 3 | 0.00 | 2.333 | 1.218 | 0.7573 | -0.0170 | 0 | 0 | |
| 4 | 0.00 | 3.333 | 1.208 | 0.6378 | -0.2059 | 0 | 0 | |
| MBD | 0 | 0.00 | 2.858 | 8.711 | 0.1561 | -0.1236 | 0 | 0 |
| Method | Seed | pSSR | Goal error (m) | Min. proxy SDF (m) | Fall | Success | ||
|---|---|---|---|---|---|---|---|---|
| MPPI | 0 | 0.00 | 0.000 | 1.999 | 0.1393 | -0.0753 | 0 | 0 |
| 1 | 0.00 | 1.250 | 1.150 | 0.2134 | -0.0792 | 0 | 0 | |
| 2 | 0.00 | 4.083 | 1.170 | 0.0737 | -0.1283 | 0 | 0 | |
| 3 | 0.00 | 1.417 | 1.235 | 0.2089 | -0.0359 | 0 | 0 | |
| 4 | 0.00 | 1.667 | 1.240 | 0.4041 | -0.0181 | 0 | 0 | |
| MBD | 0 | 0.00 | 1.708 | 4.112 | 0.1093 | -0.0400 | 0 | 0 |
| Method | Seed | pSSR | Goal error (m) | Min. proxy SDF (m) | Fall | Success | ||
|---|---|---|---|---|---|---|---|---|
| MPPI | 0 | 0.00 | 0.250 | 1.997 | 0.1389 | -0.0261 | 0 | 0 |
| 1 | 0.00 | 0.500 | 1.230 | 0.2114 | -0.0297 | 0 | 0 | |
| 2 | 0.00 | 0.917 | 1.117 | 0.0761 | -0.0681 | 0 | 0 | |
| 3 | 0.00 | 0.000 | 1.124 | 0.2080 | -0.0632 | 0 | 0 | |
| 4 | 0.00 | 0.583 | 1.280 | 0.4044 | -0.0273 | 0 | 0 | |
| MBD | 0 | 0.00 | 0.558 | 4.821 | 0.1049 | -0.0354 | 0 | 0 |
| Method | Constraint and sampling mechanism |
|---|---|
| MPPI | Gaussian action perturbations with exponential cost weighting; no geometry-aware reverse transport; 100 update iterations. |
| MBD | Model-based denoising with task cost and no explicit local constraint geometry. |
| MDOC | MBD with CBF-style derivative correction using the configured step, gain, and margin; no 2GO metric/tangent split or route/risk gate. |
| MD-COAS | Model-based diffusion with its own active-set correction, scheduled margins, and configured cap on local constraints. |
| 2GO | Reward-weighted reverse transport with task-local metric and tangent maps and the task-specific correction variants in Appendix B-C . |
| Parameter | Symbol | Value | Parameter | Symbol | Value |
|---|---|---|---|---|---|
| Reverse steps | 100 | Sampling temperature | 0.3 | ||
| Beta endpoints | Returned modes | 10 | |||
| Seeds | — | Active-row cap | 8 | ||
| CVaR tail level | 0.9 | Geometry regularizers | |||
| Base refinement gain | 0.08 | Route EMA coefficient | — | 0.7 | |
| Penalty reference | 500 | — |
| Task | Corrected candidates | |||||||
|---|---|---|---|---|---|---|---|---|
| Stones S1/S2 | 512 | 0.10 | 0.20 | 0.0 | 0.02 | — | 0 | |
| Zone A | 1024 | 0.15 | 0.15 | 0.3 | 0.02 | 0.01 | 32 | |
| Zone B | 1024 | 0.15 | 0.35 | 0.1 | 0.05 | 0.04 | 32 | |
| Zone C | 1024 | 0.15 | 0.15 | 0.3 | 0.02 | 0.01 | 32 | |
| Zone D | 1024 | 0.15 | 0.15 | 0.3 | 0.07 | 0.01 | 32 |
| Task | Term | Coefficient |
|---|---|---|
| Stepping stones | Smoothness | body step yaw foot residual rate phase rate . |
| Stepping stones | Feasibility and shape | foothold violation; support violation; geometric shape residual shape ; step-bound factors and . |
| Stepping stones | Task terms | Terminal goal weight , terminal-foot multiplier , plus progress, time guide, lateral/velocity/yaw/backward, heading, support-center, and phase-ramped terminal-foot terms. |
| Humanoid | Task terms | terminal goal, corridor centering, each arm tuck, torso yaw, height, velocity, control, progress, and squeeze term. |
| Humanoid | Zone-specific goal | in Zone C and in Zones A, B, and D. Goal deceleration is enabled only in Zone C. |
| Humanoid | Obstacle cost | ; obstacle feasibility enters through the violation evaluator, local geometry, augmented-Lagrangian schedule, and candidate correction. |
| Domain | Method | Samples | Steps | Additional parameters |
|---|---|---|---|---|
| Quadruped | MPPI | 512 | 100 | Action noise , temperature . |
| Quadruped | MBD | 512 | 100 | , extra action noise , 10 returned modes. |
| Quadruped | MDOC | 512 | 100 | , correction step , gain , margin , 10 modes. |
| Quadruped | MD-COAS | 512 | 100 | , at most 12 constraints, constraint margin , scheduler margin , 10 modes. |
| Humanoid | MPPI | 1024 | 100 | Action noise ( in Zone B), temperature . |
| Humanoid | MBD | 1024 | 100 | , extra action noise ( in Zone B), 10 modes. |