Revisiting Numerical Forecasting Models for Language-Based Trajectory Prediction
Organizations: Yonsei University · DGIST
Abstract
Language-based trajectory predictors represent coordinates as discrete tokens and learn auxiliary tasks such as destination and group reasoning. This formulation enables the model to capture behavioral intent and social context beyond coordinate dynamics alone. However, token-level objectives provide only indirect guidance for continuous coordinate-space dynamics. To address this limitation, we introduce MoRE (Mixture of Reward Experts), a refinement framework that transfers numerical forecasting priors into a pretrained language-based predictor through reinforcement learning. Five frozen numerical predictors provide complementary coordinate-level knowledge of motion and interactions. Their predictions are converted into expert rewards and combined through an uncertainty-weighted consensus that penalizes disagreement. A ground-truth reward anchors the prediction to the target trajectory. To focus refinement on difficult cases, MoRE refines the policy using the top 1% of training samples ranked by predictive entropy. Expert predictions are computed once and cached before PPO training, so the experts are not run during policy updates or inference. In this way, MoRE combines the contextual modeling of the language-based predictor with coordinate-level feedback from numerical experts. On ETH-UCY, MoRE reduces ADE from 0.22 to 0.20 m and FDE from 0.32 to 0.29 m. Relative to the base policy, ADE decreases by 17.9% on SDD and 12.7% on NBA. On ETH-UCY, MoRE also reduces collision rates and better matches ground-truth pedestrian spacing, without increasing measured inference memory or latency. The project page is available at https://jungyu0413.github.io/MoRE/.
Figures & tables
| Model | ETH | HOTEL | UNIV | ZARA1 | ZARA2 | AVG | SDD | GCS |
| Numerical models | ||||||||
| Social GAN | 0.77/1.40 | 0.43/0.88 | 0.75/1.50 | 0.35/0.69 | 0.36/0.72 | 0.53/1.04 | 13.6/24.6 | 15.9/32.6 |
| Social-STGCNN | 0.65/1.10 | 0.50/0.86 | 0.44/0.80 | 0.34/0.53 | 0.31/0.48 | 0.45/0.75 | 20.8/33.2 | 14.7/23.9 |
| PECNet | 0.61/1.07 | 0.22/0.39 | 0.34/0.56 | 0.25/0.45 | 0.19/0.33 | 0.32/0.56 | 10.0/15.9 | 17.1/29.3 |
| Trajectron++ | 0.61/1.03 | 0.20/0.28 | 0.30/0.55 | 0.24/0.41 | 0.18/0.32 | 0.31/0.52 | 11.4/20.1 | 12.8/24.2 |
| AgentFormer | 0.46/0.80 | 0.14/0.22 | 0.25/0.45 | 0.18/0.30 | 0.14/0.24 | 0.23/0.40 | 8.7/14.9 | 10.2/16.9 |
| Evaluation | Reference | MoRE (Ours) |
|---|---|---|
| Non-linear group walking | 0.297 / 0.323 | 0.211 / 0.274 |
| Non-linear mimicry walking | 0.195 / 0.277 | 0.151 / 0.217 |
| Non-linear collision avoidance | 0.251 / 0.288 | 0.198 / 0.264 |
| Non-linear subset average | 0.248 / 0.296 | 0.187 / 0.252 |
| ETH-UCY SDD | 10.03 / 17.93 | 9.14 / 15.32 |
| Inference memory (MB) | 1,401 | 1,401 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration | ADE / FDE |
|---|---|
| (a) Single-expert alignment | |
| SingularTrajectory | 0.211 / 0.315 |
| MART | 0.207 / 0.314 |
| MoFlow | 0.206 / 0.307 |
| (b) Leave-one-out removal | |
| MoRE without Social-STGCNN | 0.207 / 0.291 |
| Expert | Group | Mimicry | Non-linear | Collision | Reported avg. |
|---|---|---|---|---|---|
| None | 0.206/0.256 | 0.164/0.232 | 0.322/0.474 | 0.200/0.275 | 0.223/0.309 |
| STGCNN | 0.202/0.252 | 0.158/0.212 | 0.312/0.455 | 0.215/0.292 | 0.221/0.301 |
| DMRGCN | 0.198/0.248 | 0.148/ 0.195 | 0.318/0.462 | 0.205/0.282 | 0.224/0.305 |
| GP-Graph | 0.192 / 0.242 | 0.152/0.202 | 0.310/0.450 | 0.192 /0.268 | 0.216 /0.298 |
| SingularTraj. | 0.212/0.268 | 0.145 /0.200 | 0.292 / 0.428 | 0.202/0.278 | 0.218/0.296 |
| ExpertTraj. | 0.210/0.266 | 0.160/0.216 | 0.308/0.448 | 0.198/ 0.260 | 0.220/ 0.290 |
| Baseline | COL | Integration | COL |
|---|---|---|---|
| GP-Graph | 2.44 | Ensemble | 3.67 |
| LED | 2.62 | Knowledge distillation | 2.92 |
| MART | 2.23 | Maximum aggregation | 2.79 |
| SingularTrajectory | 2.96 | Mean aggregation | 2.37 |
| LMTraj-SUP | 2.54 | UWO (MoRE) | 2.11 |
| Method | ETH | HOTEL | UNIV | ZARA1 | ZARA2 | Reported avg. |
|---|---|---|---|---|---|---|
| Base policy | 0.409/0.501 | 0.120/0.159 | 0.218/0.344 | 0.199/0.318 | 0.175/0.272 | 0.224/0.318 |
| Supervised FT | 0.401/0.514 | 0.108/0.148 | 0.247/0.353 | 0.211/0.319 | 0.168/0.271 | 0.227/0.321 |
| Ensemble | 0.421/0.612 | 0.173/0.261 | 0.292/0.514 | 0.231/0.382 | 0.203/0.331 | 0.264/0.420 |
| Distillation | 0.382/0.492 | 0.122/0.155 | 0.211/0.338 | 0.196/0.301 | 0.177/0.277 | 0.218/0.313 |
| Reward-weighted reg. | 0.388/0.499 | 0.121/0.158 | 0.215/0.342 | 0.198/0.305 | 0.173/0.283 | 0.219/0.317 |
| Minimum-risk training | 0.412/0.515 | 0.135/0.172 | 0.230/0.355 | 0.210/0.315 | 0.188/0.298 | 0.235/0.331 |
| 0 | 0.01 | 0.05 | 0.1 | 0.5 | |
|---|---|---|---|---|---|
| 0.1 | 0.223/0.316 | 0.218/0.309 | 0.211/0.297 | 0.213/0.300 | 0.222/0.315 |
| 0.3 | 0.222/0.312 | 0.212/0.298 | 0.208/0.289 | 0.214/0.292 | 0.217/0.306 |
| 0.5 | 0.217/0.307 | 0.210/0.295 | 0.204/0.285 | 0.206/0.288 | 0.211/0.299 |
| 0.7 | 0.219/0.310 | 0.217/0.302 | 0.207/0.289 | 0.211/0.306 | 0.219/0.316 |
| 1.0 | 0.221/0.309 | 0.216/0.300 | 0.210/0.294 | 0.215/0.297 | 0.221/0.312 |
| 0 (SFT) | 0.1 | 0.3 | 0.5 | 0.7 | 1.0 | |
|---|---|---|---|---|---|---|
| ADE / FDE | 0.227/0.321 | 0.213/0.301 | 0.207/0.290 | .204/.285 | .206/.289 | .213/.303 |
| Tuning | Parameters (%) | Data (%) | ADE | FDE |
|---|---|---|---|---|
| Full FT | 100 | 1 | 0.213 | 0.297 |
| LoRA | 0.3 | 100 | 0.221 | 0.303 |
| LoRA | 0.3 | 50 | 0.220 | 0.312 |
| LoRA | 0.3 | 10 | 0.219 | 0.299 |
| LoRA | 0.3 | 0.1 | 0.223 | 0.288 |
| LoRA (Ours) | 0.3 | 1 | 0.204 | 0.285 |
| Metric | Top 1% | Random 1% | Bottom 1% | All |
| Mean entropy (nats) | 2.8 | 1.3 | 0.4 | 1.3 |
| Base-policy ADE / FDE (m) | 0.40/0.60 | 0.24/0.35 | 0.05/0.08 | 0.22/0.32 |
| Constant-velocity error (px) | 11.0 | 4.8 | 0.8 | 4.5 |
| Non-linear (%) | 82 | 45 | 2 | 41 |
| Collision-involved (%) | 25 | 13 | 1 | 9 |
| Model | ADE | FDE |
|---|---|---|
| LMTraj-SUP | 10.03 | 17.93 |
| MoRE | 9.14 | 15.32 |
| Model | Deterministic | Stochastic |
|---|---|---|
| SocialVAE | 0.54 / 1.12 | 0.21 / 0.33 |
| EigenTrajectory | 0.51 / 1.11 | 0.21 / 0.34 |
| LMTraj-SUP | 0.477 / 0.880 | 0.22 / 0.32 |
| MoRE | 0.476 / 0.877 | 0.20 / 0.29 |
| MoRE + deterministic experts | 0.477 / 0.878 | 0.20 / 0.29 |
| Method | Non-linear group walking | Non-linear mimicry walking | Non-linear collision avoidance | Subset average |
|---|---|---|---|---|
| Base policy | 0.297/0.323 | 0.195/0.277 | 0.251/0.288 | 0.248/0.296 |
| MoRE | 0.211/0.274 | 0.151/0.217 | 0.198/0.264 | 0.187/0.252 |
| Model | Memory (MB) | Training (h) | Inference (ms) |
|---|---|---|---|
| PECNet | 1,733 | 0.3 | 57.0 |
| MID | 2,929 | 6.9 | 35.0 |
| AgentFormer | 9,639 | 22.0 | 8.2 |
| SocialVAE | 1,762 | 2.1 | 73.0 |
| LMTraj-SUP | 1,401 | 3.8 | 18.3 |
| MoRE (Ours) | 1,401 | 3.8 + 1 | 18.3 |
| Platform | Latency (ms) | Rate (Hz) |
|---|---|---|
| NVIDIA DGX Spark | 55 | 18.2 |
| NVIDIA AGX Orin | 90 | 11.1 |