Diffusion models can represent complex, multimodal trajectory distributions, but extracting uncertainty from them typically requires costly Monte Carlo sampling. This limits their use in real-time control, where robots must rapidly assess risk and maintain safety margins. We introduce Score-Curvature for Online Precision Estimation (SCOPE), a lightweight module that augments diffusion trajectory models with control-ready uncertainty. SCOPE learns a structured precision matrix around each nominal trajectory by distilling score-curvature information and producing calibrated Gaussian tubes with low overhead and without repeated Monte Carlo sampling. These tubes provide per-timestep covariance estimates that can be used both as predicted occupancy for moving agents and as adaptive exploration guides for robot control. We evaluate SCOPE with mode-conditioned multimodal diffusion backbones in pedestrian forecasting, crowd navigation, Maze2D control, and real-world Franka Panda manipulation. Across these settings, SCOPE provides fast uncertainty estimation, which leads to better closed-loop performance. Project page: https://zackaxue.github.io/SCOPE-project-page/
Figures & tables
Figure 1: ( Top ) SCOPE leverages score curvature to augment diffusion predictions with precision tubes that are suitable for control. It provides uncertainty estimates at low computation cost, enabling real-time control. ( Bottom ) We apply SCOPE as a robot policy prior in MPPI and also to predict the trajectories of other agents.
Figure 2: SCOPE overview. Our method extracts control-ready Gaussian tubes from a diffusion predictor. Here, we use a multimodal diffusion backbone (left) that produces K nominal trajectories with explicit mode probabilities. The SCOPE Precision Head (right) wraps each trajectory in a structured precision matrix Λk , learned by distilling the score model’s terminal curvature via Hessian-vector products at training time. At inference, Gaussian tubes are produced with low overhead after diffusion decoding.
Figure 3: ( Left ) Y-corridor environment. ( Right ) Uncertainty quality (NLL).
Component
Time [ms]
Diffusion backbone
12.19
SCOPE Precision Head
0.60
Calibration Scaling
0.08
Total
12.87
Table 1: Inference time breakdown of SCOPE (Banded+LR).
Figure 4: ( Left ) Maze2D. ( Right ) SR and average path length vs. MPPI rollouts.
SR [%] ↑
T [s]
d [m]
Method
3
6
9
12
3
6
9
12
3
6
9
12
NF
80.0
72.0
70.0
65.0
14.44 ± 3.95
10.05 ± 6.91
8.91 ± 7.53
8.82 ± 7.63
10.14 ± 2.75
6.88 ± 4.87
5.97 ± 5.20
6.02 ± 5.37
LKF
96.0
91.0
90.0
85.0
14.64 ± 3.65
11.49 ± 6.92
10.28 ± 7.50
10.68 ± 8.87
10.26 ± 2.51
7.95 ± 4.92
6.99 ± 5.21
7.29 ± 6.06
MID-UM *
100.0
89.0
85.0
86.0
14.72 ± 3.98
11.95 ± 7.68
11.19 ± 9.59
11.84 ± 10.13
10.31 ± 2.77
8.24 ± 5.40
7.64 ± 6.64
8.09 ± 6.83
MID-MM *
100.0
91.0
85.0
88.0
14.74 ± 3.98
12.08 ± 7.47
10.56 ± 8.18
11.25 ± 9.84
10.30 ± 2.75
8.29 ± 5.22
7.21 ± 5.76
7.71 ± 6.68
MID-Ens-64
98.0
87.0
82.0
89.0
14.55 ± 3.91
11.43 ± 7.22
9.89 ± 7.72
10.59 ± 8.64
10.19 ± 2.71
7.89 ± 5.13
6.73 ± 5.41
7.24 ± 5.97
Table 2: Y-corridor navigation results across crowd densities Np∈{3,6,9,12} .
Method
SR ↑
CR ↓
TO ↓
Diffuser
0.47
0.53
0.0
Vanilla MPPI
0.17
0.80
0.03
Diff-MPPI
0.50
0.23
0.27
Ours
0.83
0.13
0.03
Table 3: Static setting.
Method
SR ↑
CR env ↓
CR human ↓
TO ↓
Ours (no WM)
0.37
0.27
0.40
0.0
MID-Ens-64
0.53
0.27
0.27
0.00
MID-Ens-128
0.47
0.30
0.33
0.00
Ours
0.70
0.13
0.10
0.07
Table 4: Dynamic setting.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Empirical structure of teacher precision matrices. The left two plots and the right two plots correspond to two representative trajectories, respectively. For each trajectory, we visualize the same precision matrix with a linear scale and a symmetric log scale. The linear-scale views show strong diagonal dominance, while the log-scale views reveal banded local structure and off-band global patterns.
Figure 6: Multimodal backbone diagnostics. (a) Effective mode count during training. Vanilla omits MI and usage regularization, MI only adds the MI reward, and Final includes both. (b) Best-of-20 displacement errors for different mode counts. (c) Three highest-probability modes in one held-out context. Thin lines show eight samples per mode, and thick lines show means over 128 samples. (d) Endpoint density contours estimated from all 128 samples per mode.
Method
Np=3
Np=6
Np=9
Np=12
LKF
0.55/1.22
0.48/1.03
0.41/0.92
0.30/0.67
MID-Ens
0.07/0.12
0.12/0.23
0.10/0.21
0.12/0.24
Ours
0.07 / 0.11
0.10 / 0.18
0.10 / 0.20
0.12 / 0.24
Appendix
Table 5: Trajectory accuracy (ADE / FDE). All SCOPE variants share the same backbone and achieve identical ADE/FDE. Bold : best.
Category
Setting
Value
Data & model
Timestep Δt
0.4 s
Observation horizon Hobs
8 steps ( 3.2 s)
Prediction horizon T
12 steps ( 4.8 s)
Discrete modes Kz / retained K
6 / 3
Diffusion steps (train / infer)
100 / 5 (DDIM)
Baselines
MID-Ens trajectory samples M
64,128,256
Appendix
Table 6: Hyperparameters for Y-corridor experiments (Q1 and Q3). Controller and cost settings apply to Q3 only; all other settings are shared.
Method
Diffusion [ms]
Head/GMM [ms]
Calibration [ms]
Total [ms]
MID-Ens ( M=64 )
26.246
0.336
–
26.581
MID-Ens ( M=128 )
47.083
0.444
–
47.527
MID-Ens ( M=256 )
91.222
0.558
–
91.779
Banded+LR (Ours)
12.19
0.60
0.08
12.87
Appendix
Table 7: Inference time breakdown ( Np=3 ): MID-Ensemble at M∈{64,128,256} samples versus Banded+LR (Ours). All diffusion passes are fully batched on a single GPU. Ours shares the same backbone but avoids repeated sampling, yielding consistently lower total latency.
Figure 7: Coverage diagnostic on the 3-pedestrian Y-junction. NLL calibration improves empirical coverage relative to the uncalibrated tubes, but is not coverage-targeted and does not provide formal coverage guarantees.
Diffusion trajectory prior (Diff-MPPI and Ours)
Backbone
1D temporal U-Net
dim 64 , mults (1,4,8)
Wall encoder
1D vision transformer
embed 64 , depth 3
Horizon H
48
Diffusion steps (train / infer)
100 / 5 (DDIM)
CFG dropout / sampling weight
0.2 / 2.0
Training steps
2×106
Appendix
Table 8: Hyperparameters for the Maze2D experiment (Q2). The diffusion prior is shared by Diff-MPPI and Ours; the Precision Tube Head is exclusive to Ours; the MPPI block applies to all four methods.
Planner
SR ↑
CR ↓
GR ↑
L [m] ↓
Diffuser
0.874
0.126
1.0
6.134
PB-Diff
0.887
0.113
1.0
6.042
Vanilla MPPI
0.9375
0.0125
0.9500
5.573 ± 3.354
Log-MPPI
0.9455
0.0065
0.9510
5.994 ± 4.506
Diff-MPPI
0.9635
0.0260
0.9875
5.328 ± 1.601
SCOPE (Ours)
0.9925
0.0065
0.9990
4.996 ± 0.763
Appendix
Table 9: Maze2D navigation results. MPPI-based planners use N=1024 rollouts per planning step.
Figure 8: Qualitative comparison of uncertainty tubes from Direct NLL PH (top) and SCOPE (bottom). Columns show selected examples with 3, 6, 9, and 12 pedestrians. Colors identify predicted modes. Shaded ellipses visualize uncertainty along each mode. Dashed gray lines show observed histories, and solid black lines show ground-truth futures. Each column uses matching axis limits across rows. Several Direct NLL PH tubes extend far beyond their nominal trajectories, whereas SCOPE produces more localized uncertainty in these cases.
SR [%] ↑
Method
3 pedestrians
6 pedestrians
9 pedestrians
12 pedestrians
Direct NLL PH
100
91
84
81
SCOPE (Ours)
100
96
93
92
Appendix
Table 10: Navigation success rate across pedestrian densities. Bold indicates the best result for each density.
Figure 9: Real-robot shelf setup for pick-and-place task. The Franka arm operates in a constrained shelf workspace. In this example, the OREO box is the human target and the tomato is the robot target, while the remaining objects serve as static obstacles for collision-aware MPPI control.
Figure 10: Real-robot setup and representative approach modes. (A) Front approach, where the arm reaches the target through a wider opening. (B) Side approach, where the arm reaches the target with tighter clearance. (C) Low-cost RGB webcam setup used for hand tracking in the dynamic experiments.
Figure 11: Representative real-robot failure cases. A. Static Diff-MPPI failure. Without SCOPE, the planner lacks clearance- and flexibility-aware mode selection and may choose the tighter side approach. Near the object, a potential shelf collision blocks further progress; limited-horizon MPPI cannot retreat and reorient, so the trial times out (TO). B. Dynamic failure of Ours (no WM), where SCOPE guides the robot policy but is not used as a human-motion world model. Without hand prediction, the controller reacts only after the hand enters the shelf workspace, causing a late upward evasive motion that collides with the shelf.