Masked Generative Motion Planning with Geometry-Guided Token Search
Organizations: University of Glasgow, Glasgow, United Kingdom
Abstract
Generative motion planners typically use learned trajectory priors for initial generation, while leaving test-time repair to local continuous refinement. We introduce Masked Generative Motion Planning (MGMP), which extends the learned prior from efficient parallel generation to structural repair. A masked generative transformer generates discrete trajectory candidates in parallel, and Geometry-Guided Token Search (GGTS) uses scene geometry to target where to edit and which prior-supported alternatives to evaluate. This turns refinement into an efficient search over discrete motion alternatives, enabling route-level restructuring beyond local trajectory deformation. MGMP achieves 96% success on Ring Maze and 82% repair success on Controlled Route Invalidation on Kuka, exceeding the strongest external baselines by 23 and 25 percentage points, respectively. It further generalizes to unseen layouts, additional obstacles, unseen geometries, single- and dual-arm planning, and real-world Baxter tasks.
Figures & tables
| Methods | Ring Maze | Controlled Route Invalidation on Kuka | |||||
|---|---|---|---|---|---|---|---|
| SR | VSF | RT | RS | Avg. Dev. | Max. Dev. | Changed Span | |
| CHOMP | ±0.09 | ±0.10 | ±0.14 | ±0.15 | ±0.19 | ±0.32 | ±9.40 |
| TrajOpt | ±0.04 | ±0.03 | ±0.01 | ±0.14 | ±0.20 | ±0.38 | ±10.10 |
| STOMP | ±0.01 | ±0.03 | ±0.07 | ±0.15 | ±0.22 | ±0.35 | ±10.80 |
| MPD | ±0.06 | ±0.02 | ±0.06 | ±0.13 | ±0.20 | ±0.35 | ±9.60 |
| PBD | ±0.10 | ±0.09 | ±0.05 | ±0.14 | ±0.22 | ±0.39 | ±10.30 |
| Methods | Maze2D-6 | Concave Maze2D | 7-DoF Kuka | 14-DoF DualKuka | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SR | VSF | PT | SR | VSF | PT | SR | VSF | PT | SR | VSF | PT | |
| RMP | ±0.13 | – | ±0.59 | ±0.13 | – | ±0.64 | ±0.14 | – | ±0.72 | ±0.18 | – | ±0.88 |
| RRT* | ±0.01 | – | ±0.13 | ±0.03 | – | ±0.22 | ±0.17 | – | ±0.78 | ±0.17 | – | ±0.64 |
| P-RRT* | ±0.02 | – | ±0.18 | ±0.05 | – | ±0.27 | ±0.18 | – | ±0.80 | ±0.17 | – | ±0.64 |
| BIT* | ±0.01 | – | ±0.02 | ±0.01 | – | ±0.04 | ±0.02 | – | ±0.25 | ±0.14 | – | ±0.76 |
| CHOMP | ±0.01 | ±0.07 | ±0.03 | ±0.01 | ±0.02 | ±0.08 | ±0.04 | ±0.02 | ±0.08 | ±0.05 | ±0.04 | ±0.12 |
| Methods | Kuka Add.-Obs. | DualKuka Add.-Obs. | Kuka Unseen Geo. | DualKuka Unseen Geo. | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SR | VSF | PT | SR | VSF | PT | SR | VSF | PT | SR | VSF | PT | |
| RMP | ±0.22 | – | ±1.12 | ±0.19 | – | ±0.88 | ±0.15 | – | ±0.78 | ±0.17 | – | ±0.80 |
| RRT* | ±0.23 | – | ±1.03 | ±0.16 | – | ±0.56 | ±0.15 | – | ±0.73 | ±0.19 | – | ±0.74 |
| P-RRT* | ±0.23 | – | ±1.02 | ±0.17 | – | ±0.54 | ±0.15 | – | ±0.70 | ±0.18 | – | ±0.70 |
| BIT* | ±0.10 | – | ±0.62 | ±0.21 | – | ±0.81 | ±0.10 | – | ±0.58 | ±0.19 | – | ±0.96 |
| CHOMP | ±0.02 | ±0.04 | ±0.05 | ±0.07 | ±0.06 | ±0.01 | ±0.02 | ±0.01 | ±0.06 | ±0.08 | ±0.05 | ±0.08 |
| Variant | Ring Maze | 7-DoF Kuka | ||||
|---|---|---|---|---|---|---|
| SR | VSF | RT | SR | VSF | RT | |
| MGT (Conf.-Refine) | 0.69 | 0.20 | 0.03 | 0.84 | 0.58 | 0.07 |
| MGT + EmbOpt-30 | 0.75 | 0.29 | 0.08 | 0.90 | 0.74 | 0.30 |
| Conf.-only Remask | 0.84 | 0.49 | 0.08 | 0.92 | 0.78 | 0.43 |
| Geom.-only Remask | 0.93 | 0.64 | 0.08 | 0.95 | 0.84 | 0.44 |
| Prob.-only Screening | 0.88 | 0.55 | 0.07 | 0.94 | 0.81 | 0.41 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Maze2D | 7-DoF Kuka | 14-DoF DualKuka |
|---|---|---|---|
| Codebook size | 8192 | 8192 | 8192 |
| Code dimension | 16 | 256 | 256 |
| Encoder width | 16 | 256 | 256 |
| Encoder depth | 4 | 4 | 4 |
| Downsampling rate | 2 | 2 | 1 |
| Stride size | 2 | 2 | 2 |
| Parameter | Maze2D | 7-DoF Kuka | 14-DoF DualKuka |
| GGTS refinement rounds | 2 | 2 | 2 |
| Token updates per round | 2 | 2 | 2 |
| Proposed code candidates per update | 64 | 16 | 16 |
| Candidates retained after screening | 32 | 2 | 2 |
| Maximum verification remask ratio | 0.4 | 0.4 | 0.4 |
| Polishing iterations | 5 | 5 | 5 |
| (a) CHOMP: optimization iterations | ||||
|---|---|---|---|---|
| Iterations | 10 | 20 | 30 | 50 |
| SR (%) | 73.30 | 73.30 | 73.30 | 73.30 |
| VSF (%) | 22.30 | 23.30 | 23.70 | 23.30 |
| Variant | Position Selection | Candidate Shortlist | Exact Verification | Commit | Continuous Opt. |
|---|---|---|---|---|---|
| MGT (Conf.-Refine) | Confidence | MGT prob. | – | – | – |
| MGT + EmbOpt-30 | – | – | – | – | EmbOpt-30 |
| Conf.-only Remask | Confidence | Eq. (2), | top- | Sequential | Polish |
| Geom.-only Remask | Geometry | Eq. (2), | top- | Sequential | Polish |
| Probability-only Screening | Geom.+Conf. | MGT prob., | top- | Sequential | Polish |
| Independent Exact Search | Geom.+Conf. | Eq. (2), | top- | Independent | Polish |
| (a) Candidate shortlist size | (b) Refinement schedule | (c) Continuous polishing | |||||||||
| SR | VSF | RT | Schedule | SR | VSF | RT | Steps | SR | VSF | RT | |
| 1 | 0.95 | 0.84 | 0.38 | 0.93 | 0.79 | 0.34 | 0 | 0.94 | 0.83 | 0.40 | |
| 2 | 0.97 | 0.87 | 0.44 | 0.95 | 0.83 | 0.40 | 3 | 0.96 | 0.86 | 0.42 | |
| 4 | 0.97 | 0.88 | 0.56 | 0.97 | 0.87 | 0.44 | 5 | 0.97 | 0.87 | 0.44 | |
| 16 | 0.98 | 0.89 | 1.28 | 0.96 | 0.85 | 0.52 | 10 | 0.97 | 0.88 | 0.48 | |