MoCAR: Motion-code Coordinate-aware AutoRegression for Continuous Trajectory Forecasting
Authors: Yiming Xu, Hao Cheng, Monika Sester
Organizations: Institute of Cartography and Geoinformatics Leibniz University Hannover, 30167 Hannover, Germany · Faculty of Geo-Information Science and Earth Observation University of Twente, 7522 NH Enschede, The Netherlands
Autoregressive generation is natural for language, where predicted tokens can be directly reused as the next prediction state, but trajectory forecasting lacks such a clean token: motion is continuous, multimodal, and expressed in local coordinate frames that evolve with the predicted trajectory. We present MoCAR (Motion-code Coordinate-aware AutoRegression), a decoder-only framework that casts trajectory forecasting as next-code prediction in a coordinate-aware continuous latent space. MoCAR learns a continuous motion-code space from endpoint-normalized trajectory segments, where each code jointly captures local trajectory geometry and the reference-frame transition induced by that segment. Historical motion codes are used as a teacher-forced prefix, future codes are generated autoregressively under temporal, map, agent, and mode interactions, and predicted codes persist in latent memory while decoded endpoints update the local scene context. This enables rollout without trajectory-space re-tokenization, trajectory queries, goal candidates, or proposal-and-refinement pipelines. On Argoverse (AV) benchmarks, MoCAR achieves top-tier performance with a simple single-stage architecture, transfers strongly from AV2 to AV1 in zero-shot evaluation, and improves on turn-heavy scenarios. Ablations confirm that the learned continuous motion-code space, latent alignment, weak KL regularization, and joint tokenizer-predictor optimization are essential for stable latent autoregression.
Figures & tables
Figure 1: Paradigm comparison. 1 Query-based encoder–decoder methods generate trajectory proposals and refine them through additional decoding stages. 2 Trajectory-space autoregressive methods explicitly decode each sub-trajectory, update the local frame, and re-tokenize the predicted motion for the next step. 3 In contrast, MoCAR represents endpoint-normalized trajectory segments as continuous motion codes and performs autoregressive next-code prediction directly in latent space, avoiding explicit trajectory queries and proposal-and-refinement pipelines.
Figure 2: Coordinate-aware motion segment. Each local trajectory segment is normalized in the coordinate frame of its endpoint anchor and encoded as a continuous motion code. This code jointly represents local geometry and the induced frame transition, making it the autoregressive prediction unit in MoCAR.
Figure 3: Decoder-only motion-code autoregression in MoCAR. Historical segments are tokenized into a teacher-forced motion-code prefix, and future prediction rolls out autoregressively in latent space. Mode embeddings separate latent rollout branches, while the predictor updates each branch using temporal, map, agent, and mode interactions before decoding the next segment and endpoint anchor. MoCAR therefore forecasts directly in code space, without trajectory-space re-tokenization or proposal-and-refinement pipelines.
Method
minFDE 1 ↓
minADE 1 ↓
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
FRM ( Park et al., 2023 )
5.93
2.37
1.81
0.89
0.29
2.47
HDGT ( Jia et al., 2023 )
5.37
2.08
1.60
0.84
0.21
2.24
MTR ( Shi et al., 2022 )
4.39
1.74
1.44
0.73
0.15
1.98
HPTR ( Zhang et al., 2023 )
4.61
1.84
1.43
0.73
0.19
2.03
HeteroGCN ( Gao et al., 2023 )
4.40
1.72
1.34
0.69
0.18
1.90
ProphNet ( Wang et al., 2023 )
4.74
1.80
1.33
0.68
0.18
1.88
Table 1: Comparison with state-of-the-art non-ensemble methods on the AV2 test set. MoCAR reaches DeMo/DONUT-level performance without trajectory queries or explicit refinement. Best results are in bold, and second-best results are underlined.
Method (AV2 → AV1, zero-shot)
minFDE 1 ↓
minADE 1 ↓
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
DeMo ( Zhang et al., 2024 )
5.70
2.52
2.03
1.08
0.25
2.74
QCNet ( Zhou et al., 2023 )
5.60
2.51
1.55
0.95
0.19
2.27
DONUT ( Knoche et al., 2025 )
5.05
2.32
1.58
0.97
0.19
2.26
MoCAR (Ours)
5.14
2.37
1.24
0.72
0.11
1.93
Table 2: Zero-shot transfer to AV1 validation. All methods use AV2-trained checkpoints with identical preprocessing. Full AV1 evaluation appears in Appendix F .
Method
Subset
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
QCNet ( Zhou et al., 2023 )
Turn-heavy ( ≥45∘ )
2.37
1.17
0.36
2.98
Lower-turning ( <45∘ )
1.10
0.66
0.13
1.72
All
1.25
0.72
0.16
1.87
DONUT ( Knoche et al., 2025 )
Turn-heavy ( ≥45∘ )
2.01
1.10
0.31
2.64
Lower-turning ( <45∘ )
1.07
0.67
0.12
1.70
All
1.18
0.73
0.14
1.81
Table 3: Performance on the AV2 validation set stratified by future turning angle. The turn-heavy subset contains samples whose ground-truth future trajectory has a heading change of at least 45∘ .
Method
Stride
#Params
GPU inference (ms)
QCNet ( Zhou et al., 2023 )
–
7.7M
35.6
DONUT ( Knoche et al., 2025 )
10
9.0M
126.7
MoCAR (matched stride)
10
6.3M
97.9
MoCAR (default)
15
6.3M
66.8
Table 4: Parameter count and GPU inference latency on a single Nvidia RTX 4090, measured per scenario with batch size B=1 . Stride is the number of timesteps advanced per segment and does not apply to QCNet.
Figure 4: Learned motion-code space. t-SNE visualization of encoder means from 50,000 AV2 validation segments, colored by agent type, mean speed, and turn angle. The learned space organizes segments by agent category and motion dynamics, suggesting that the latent codes capture fine-grained motion primitives.
Variant
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
Code dim
z dim = 16
1.23
0.72
0.153
1.88
z dim = 24
1.17
0.70
0.134
1.84
z dim = 32
1.18
0.70
0.138
1.85
Stride
stride = 10
1.18
0.70
0.133
1.84
stride = 15
1.17
0.70
0.134
1.84
stride = 20
1.18
0.70
0.138
1.85
Table 5: Ablation on motion-code design on the AV2 validation set.
Variant
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
Raw geometric MLP encoder–decoder
1.29
0.73
0.17
1.92
Same-architecture GRU (forecasting only)
1.27
0.73
0.16
1.91
w/o latent-alignment loss Lalign
1.23
0.72
0.15
1.89
Autoencoder tokenizer ( βkl=0 )
1.29
0.75
0.16
1.95
Frozen pretrained tokenizer–de-tokenizer
1.25
0.73
0.15
1.90
MoCAR (re-encoding feedback)
1.21
0.72
0.15
1.87
Table 6: Ablation on motion-code representation, training, and feedback on AV2 validation. Best results are in bold.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Term
Value
VAE objective λvae
0.1
KL coefficient βkl
10−4
Future-trajectory L1 (winner) λtraj
5.0
Historical-trajectory L1 λtrajhist
1.0
Heading-trajectory L1 (hist / fut)
1.0 / 1.0
Latent alignment λalign
0.01
Appendix
Table 7: Loss weight settings.
Term
Weight
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
KL βkl
0
1.29
0.75
0.16
1.95
10−4 (default)
1.17
0.70
0.13
1.84
Alignment λalign
0
1.23
0.72
0.15
1.89
0.01 (default)
1.17
0.70
0.13
1.84
1.0
1.19
0.71
0.14
1.86
Appendix
Table 8: Representation-loss weight comparisons on the AV2 validation set. Each group changes one coefficient while keeping the others at their defaults. The zero-weight and default results are repeated from Table 6 .
Variant
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
3 context-update blocks
1.27
0.74
0.156
1.91
4 context-update blocks
1.22
0.72
0.142
1.89
5 context-update blocks (MoCAR)
1.17
0.70
0.134
1.84
6 context-update blocks
1.17
0.70
0.133
1.84
Appendix
Table 9: Predictor depth sweep on the AV2 validation set.
AV2 validation
AV1 zero-shot
Anchor
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
minFDE 6 ↓
Segment start
1.19
0.71
0.14
1.85
1.25
Segment endpoint
1.17
0.70
0.13
1.84
1.24
Appendix
Table 10: Choice of segment anchor with local coordinate normalization. Both variants are trained on AV2 and evaluated on AV2 validation and AV1 validation without fine-tuning.
Feedback state
1.5 s
3.0 s
4.5 s
6.0 s
Trajectory: decode and re-encode
0.19
0.44
0.80
1.21
Persistent motion code (MoCAR)
0.20
0.44
0.78
1.17
Appendix
Table 11: Autoregressive feedback ablation on AV2 validation, reporting minFDE 6 by prediction horizon.
k=1
2
3
4
5
6
Entropy H
MoCAR
0.177
0.134
0.119
0.172
0.200
0.199
1.774
Appendix
Table 12: Winner-mode distribution and entropy on the AV2 validation set.
Motion type
#Segments
Pos. err. (m) ↓
Head. err. (rad) ↓
Stationary ( v<0.1m/s )
397488
0.02
0.00
Straight ( 0∘≤∣Δθ∣<5∘ )
238393
0.08
0.01
Mild turn ( 5∘≤∣Δθ∣<22.5∘ )
33713
0.10
0.03
Sharp turn ( 22.5∘≤∣Δθ∣<45∘ )
11514
0.12
0.06
Very sharp turn (turn-heavy) ( ∣Δθ∣≥45∘ )
860
0.16
0.19
Appendix
Table 13: VAE reconstruction error on AV2 validation segments, grouped by motion type. Stationary segments are grouped by mean speed, while non-stationary segments are grouped by segment-level heading change. Position error is measured by the mean L2 distance over segment points, and heading error is measured by the mean absolute wrapped angle error. Turn categories are defined over local P=16 segments, about 1.5s on AV2; thus ∣Δθ∣≥22.5∘ already corresponds to a sharp local turn.
Method
Stage
Subset
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
QCNet ( Zhou et al., 2023 )
w/o refinement
Turn-heavy ( ≥45∘ )
2.40
1.17
0.38
2.99
Lower-turning ( <45∘ )
1.13
0.65
0.14
1.73
All
1.28
0.72
0.17
1.88
w/ refinement
Turn-heavy ( ≥45∘ )
2.37
1.17
0.36
2.98
Lower-turning ( <45∘ )
1.10
0.66
0.13
1.72
All
1.25
0.72
0.16
1.87
Appendix
Table 14: Effect of method-specific refinement stages on the AV2 validation set stratified by future turning angle. All methods are evaluated on the same 24,988 candidates and the same ground-truth turn split as Table 3 , with 3,007 turn-heavy and 21,981 lower-turning samples. Best values are in bold, and second-best values are underlined; ties are marked consistently.
Property
AV2 training
AV1 zero-shot evaluation
Observation / prediction
5 s / 6 s
2 s / 3 s
Available motion states
Position, velocity, heading
Position; velocity and heading derived
Agent labels
Native AV2 categories
Ego and AGENT mapped to vehicle; others to other/unknown
Model weights
Trained on AV2
Unchanged AV2 checkpoint; no fine-tuning
Sampling / code duration
10 Hz / 1.5 s
10 Hz / 1.5 s
Appendix
Table 15: AV2 training and AV1 zero-shot evaluation settings. A segment contains 16 points and spans 1.5 s at the shared sampling rate.
Noise augmentation
minFDE 1 ↓
minADE 1 ↓
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
b-minFDE 6 ↓
Without
5.21
2.40
1.27
0.73
0.11
1.97
With (default)
5.14
2.37
1.24
0.72
0.11
1.93
Appendix
Table 16: Effect of training-time positional-noise augmentation on AV2-to-AV1 zero-shot evaluation. Both models use the same AV1 preprocessing and are evaluated without noise augmentation.
Method
minFDE 6 ↓
minADE 6 ↓
MR 6 ↓
TNT ( Zhao et al., 2021 )
1.29
0.73
0.09
LaneRCNN ( Zeng et al., 2021 )
1.19
0.77
0.08
mmTransformer ( Liu et al., 2021 )
1.15
0.71
0.11
LaneGCN ( Liang et al., 2020 )
1.08
0.71
–
DenseTNT ( Gu et al., 2021 )
1.05
0.73
0.10
SSL-Lanes ( Bhattacharyya et al., 2023 )
1.01
0.70
0.09
Appendix
Table 17: Full Argoverse 1 validation evaluation. Top block: representative methods originally trained on AV1 (numbers from their original papers). Middle block: zero-shot transfer of AV2-trained models, repeated from Table 2 for completeness. Bottom block: MoCAR under fine-tuning and full AV1 training.
Figure 5: Qualitative comparison on challenging AV2 scenarios. From left to right, we show predictions from QCNet, DONUT, MoCAR, and the ground truth. Blue denotes the observed history, orange denotes predicted future trajectories, and green denotes the ground-truth future. While all methods are generally map-aligned, MoCAR produces more coherent and temporally consistent predictions in complex intersections and interaction-heavy scenes.
Figure 6: Additional qualitative success cases on the AV2 validation set. Eight representative examples covering roundabouts, straight driving, curved-road following, and turning maneuvers at intersections. Blue shows the observed history, green the ground-truth future, and orange the predicted futures, with darker shades indicating higher predicted probability. The predictions remain consistent with the local map structure and closely follow the ground-truth motion.
Figure 7: Representative failure cases on the AV2 validation set. Blue shows the observed history, green the ground-truth future, and orange the predicted futures, with darker shades indicating higher predicted probability. Most errors occur when the model follows the most map-consistent continuation, but the realized behavior reflects a later or less common maneuver.
Motion forecasting often requires trading interpretability for predictive accuracy. Standard anchor-based architectures rely on opaque latent queries that are highly prone to latent collapse, or naive trajectory sampling that limits multi-modal diversity. We propose an end-to-end differentiable framework that grounds predictions in a comprehensive "motion bank", a structured embedding space of physically realizable trajectories constructed via contrastive learning. Rather than regressing paths from a blank slate, our architecture dynamically retrieves explicit motion priors using a novel Anchor Retrieval Layer. This module adapts orthogonally initialized queries via a Dual-Level Gated Cross-Attention mechanism and executes discrete trajectory selection using a Straight-Through Gumbel-Softmax estimator to preserve continuous gradient flow. The retrieved semantically grounded anchors are then geometrically refined by a DETR-style decoder, optimized jointly with a Winner-Takes-All (WTA) kinematic Gaussian Mixture Model (GMM), a latent diversity penalty, and a soft-min weighted endpoint loss. By strictly conditioning the decoding phase on diverse, interpretable motion primitives, our approach eliminates the "black box" of standard latent queries while achieving competitive multi-modal accuracy on the Argoverse 2 and Waymo Open Motion datasets. Code is available at: https://github.com/abviv/recall2predict
Abhishek Vivekanandan, Ahmed Abouelazm, J. Marius Zöllner
FZI Forschungszentrum Informatik · Karlsruhe Institute of Technology (KIT) Karlsruhe, Germany
Language-based trajectory predictors represent coordinates as discrete tokens and learn auxiliary tasks such as destination and group reasoning. This formulation enables the model to capture behavioral intent and social context beyond coordinate dynamics alone. However, token-level objectives provide only indirect guidance for continuous coordinate-space dynamics. To address this limitation, we introduce MoRE (Mixture of Reward Experts), a refinement framework that transfers numerical forecasting priors into a pretrained language-based predictor through reinforcement learning. Five frozen numerical predictors provide complementary coordinate-level knowledge of motion and interactions. Their predictions are converted into expert rewards and combined through an uncertainty-weighted consensus that penalizes disagreement. A ground-truth reward anchors the prediction to the target trajectory. To focus refinement on difficult cases, MoRE refines the policy using the top 1% of training samples ranked by predictive entropy. Expert predictions are computed once and cached before PPO training, so the experts are not run during policy updates or inference. In this way, MoRE combines the contextual modeling of the language-based predictor with coordinate-level feedback from numerical experts. On ETH-UCY, MoRE reduces ADE from 0.22 to 0.20 m and FDE from 0.32 to 0.29 m. Relative to the base policy, ADE decreases by 17.9% on SDD and 12.7% on NBA. On ETH-UCY, MoRE also reduces collision rates and better matches ground-truth pedestrian spacing, without increasing measured inference memory or latency. The project page is available at https://jungyu0413.github.io/MoRE/.
Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavior and plan safe maneuvers. We observe that current state-of-the-art forecasting models on Argoverse 2 and the Waymo Open Motion Dataset tailor their training objectives to the different benchmark metrics. Because these metrics encourage conflicting behavior, we propose a paradigm change for trajectory forecasting: training models with metric-agnostic probabilistic objectives and treating metric optimization as a downstream task applied to the predictive distribution. Concretely, we introduce Trajectory Distribution Evaluation (TraDiE) policies, metric-specific policies that map a predictive distribution to the set of K trajectories and confidences required by trajectory forecasting metrics. We evaluate this framework by introducing DONUT-NLL, which adapts the training objective of the state-of-the-art trajectory forecasting model DONUT to directly optimize the predictive distribution. Using our policies, DONUT-NLL achieves state-of-the-art results on all metrics of the Waymo motion prediction benchmark.
Markus Knoche, Daan de Geus, Bastian Leibe
RWTH Aachen University, Germany · Eindhoven University of Technology, Netherlands