Multimodal imitation learning requires diverse executable futures under the same observation and consistent behavior across replanning cycles. We present Conditional Trajectory Peaks (CTP), a single-pass policy framework that jointly predicts complete action-chunk candidates, probability masses, and trajectory scales. Distribution-Aware Peak Specialization (DAPS) specializes trajectory peaks using trajectory-level posterior responsibilities and mass- and scale-modulated overlap constraints. Evidence-Gated Trajectory Belief Transport (ETBT) maintains cross-chunk consistency through geometric correspondence between exchangeable candidate sets, while allowing current policy evidence to override historical constraints. CTP achieves a coverage score of 91.40% on Push-T; success rates of 100.0%, 79.72%, and 84.44% on D3IL Avoiding, Aligning, and Sorting-2, respectively. On LIBERO, CTP achieves an average success rate of 97.25%. In real-world dual-arm experiments, CTP preserves both placement modes in a two-plate task, succeeding in all 50 trials. On bottle uprighting and pen placement into a holder, it maintains success rates comparable to π0.5 while reducing policy inference latency from 218.24 ms to 75.80 ms. These results demonstrate that single-pass trajectory modeling can combine multimodal behavior, closed-loop consistency, and efficient inference.
Figures & tables
Fig. 1: Core challenges in multimodal closed-loop manipulation. (a) The same observation admits multiple feasible futures, while the conditional mean may be infeasible. (b) Diffusion and flow matching policies generate action chunks iteratively, increasing online decision latency. (c) Naive single-pass multi-candidate prediction can still suffer from mode collapse during training and inconsistent execution caused by mode switching during receding-horizon replanning.
Fig. 2: Overview of CTP. (a) Training: the observation history passes through a shared policy backbone and trajectory head once to generate K trajectory peaks with probability masses and uncertainty. DAPS promotes mode specialization through trajectory-level responsibility assignment and inter-peak overlap constraints. (b) Inference: ETBT propagates historical behavioral beliefs through geometric correspondences between candidate trajectories in successive replanning cycles. Current policy evidence determines whether historical constraints are retained or released.
Method
Reported results
Unified evaluation
IBC/DFO [ 2 , 8 ]
0.90/0.84
0.8340
BeT [ 2 , 9 ]
0.79/0.70
0.7817
Energy Policy [ 12 ]
0.85
0.7300
Diffusion Policy-C [ 2 ]
0.95/0.91
0.8790
IMLE Policy [ 6 ]
0.59/0.54
0.8874
Liquid-MDN [ 13 ]
0.91
0.7394
TABLE I: Closed-loop performance of multimodal policies on Push-T.
Method
Params. (M)
NFE
Mean (ms)
P95 (ms)
CTP
74.041
1
0.9507
0.9745
Diffusion Policy
65.783
100
902.8522
904.8401
TABLE II: Model size and policy invocation latency on the same hardware.
Task
BC-MLP
DDPM-MLP
Best prior result
CTP
Avoiding
66.6±51.2
63.7±5.5
BESO 95.0±1.5
100.0
Aligning
70.8±5.2
76.3±3.9
VAE-ACT 89.1±2.2
79.72
Sorting-2
44.4±6.9
46.0±3.9
DDPM-ACT 88.2±2.3
84.44
TABLE III: Closed-loop success rates (%) on D3IL multimodal control tasks.
Fig. 3: Qualitative comparison on D3IL Avoiding. The deterministic MLP trained with MSE (top) and full CTP (bottom) start from the same initial state. Each row presents eight frames spanning a complete rollout; the bird’s-eye view (BEV) shows the corresponding executed trajectories. In this example, the MLP collides with an obstacle, while CTP successfully reaches the goal.
Method
Spatial
Object
Goal
Long
Average
OpenVLA [ 21 ]
84.7
88.4
79.2
53.7
76.50
π0 [ 22 ]
96.8
98.8
95.8
85.2
94.15
LingBot-VA [ 23 ]
98.5
99.6
97.2
98.5
98.45
Motus [ 24 ]
96.8
99.8
96.6
97.6
97.70
Fast-WAM [ 25 ]
98.2
100.0
97.0
95.2
97.60
π0.5 [ 5 ]
97.0
99.0
98.0
93.0
96.75
TABLE IV: Comparison of results on LIBERO.
Fig. 4: Success–latency trade-off on LIBERO. CTP achieves an average success rate of 97.25% with a mean policy-call latency of 75.80 ms. All policy-call latencies were measured locally at batch size 1 on identical hardware comprising an NVIDIA GeForce RTX 4090 GPU and two Intel Xeon Gold 6530 CPUs. Success rates for published baselines are taken from their respective papers, whereas CTP and π0.5 are evaluated under the same local protocol.
Fig. 5: Placement into either of two plates. (a) The real-world setup. After grasping the target object, the robot may place it in either the left or right plate; both outcomes count as success. (b) Demonstration trajectories visualized in a 3D scene model. Orange and cyan trajectories correspond to the left- and right-plate placement modes, respectively. We collected 300 demonstration episodes, with 150 per mode, yielding a balanced, spatially separated bimodal trajectory distribution.
Fig. 6: Real-world bottle uprighting and pen placement experiments and comparative results. (a) Dual-arm robot platform and camera configuration. (b) Representative execution sequences, task success rates, and per-call policy inference latencies of π0.5 and CTP on the same NVIDIA RTX 4090 GPU.
Configuration
Mean Score (%)
Single-peak baseline ( K=1 )
77.34
Multi-peak baseline ( K=4 )
80.30
Multi-peak + DAPS
85.67
CTP (full)
91.40
TABLE V: Incremental component ablations of CTP on Push-T, measured by Mean Score (%).
Robotic manipulation requires the effective integration of heterogeneous inputs, including visual observations, language instructions, and trajectory representations, to generate accurate actions. Existing transformer-based policies typically process these heterogeneous modalities within a shared parameter space, which often leads to modality interference and inefficient representation learning, especially in data-scarce scenarios. While Mixture-of-Experts (MoE) offers a scalable solution through expert specialization, conventional routing mechanisms are often sensitive to such cross-modal representation discrepancies, resulting in unstable expert assignment and expert collapse. In this work, we propose MATE (Multi-ModAl TrajEctory Policies), a novel trajectory prediction framework built upon MoE. Specifically, we introduce a Multi-Modal MoE architecture to achieve fine-grained sub-token feature decoupling, and design a cross-modal cosine router for stable and scale-invariant expert assignment across heterogeneous modalities. We further employ temperature-controlled routing and stochastic noise injection to improve expert balance and prevent premature routing collapse under scarce demonstrations. Experiments on the LIBERO benchmark show that our MATE consistently outperforms prior work under data scarcity. It achieves a 4.75% improvement in average success rate over the trajectory-guided counterpart. Real-world experiments on robotic ping-pong also suggest that the predicted trajectories can provide useful guidance for downstream robotic execution, further indicating the practical feasibility of our algorithm.
Zijia Chen, Yuenan Hou, Xinhua Jiang +3
College of Electronic Science and Technology, National University of Defense Technology Changsha, 410073, China · Shanghai AI Laboratory Shanghai, 200000, China
Behavioral cloning becomes challenging when the same observation admits several valid actions. We study how generative behavioral-cloning policies represent such multimodal expert behavior and identify different bottlenecks across model parameterizations. For latent-variable policies, preserving demonstrated modes requires action-conditioned information in the latent representation. Excessive posterior-prior regularization can suppress this information and prevent the policy from distinguishing demonstrated modes. Weaker or aggregate regularization can preserve mode information, but shifts the challenge to ensuring that the deployment-time prior covers the relevant latent regions. For action-space generative policies, multimodality is constrained by the smoothness of the base-to-action transport: a map with a small Lipschitz constant cannot assign substantial probability to many well-separated modes. Covering many modes therefore requires either sharp transitions in base space or off-support bridge regions in action space. Experiments on synthetic multimodal navigation and a physical-robot bimodal manipulation task support these mechanisms. In contrast, our analysis reveals limited conditional multimodality in standard robotic simulation benchmarks, where deterministic regression remains competitive.
Lorenzo Mazza, Massimiliano Datres, Ariel Rodriguez +3
NCT/UCC Dresden, UKDD Dresden, TU Dresden, DKFZ Heidelberg · Ludwig-Maximilians-Universität München, Munich Center for Machine Learning (MCML) · BMFTR Research Hub 6G-Life +4
Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time-critical settings like autonomous driving. Fast-sampling variants using DDIM and informed initial noise distributions partially alleviate this issue, but they either fail to achieve true single-step generation or are constrained by the chosen noise distribution. Consistency Models (CMs) offer high-quality one-step generation by mapping noise directly to data, but are difficult to train from scratch. We propose ECTraj, an enhanced CM pipeline with improved training and conditional generation for trajectory prediction. Our framework extends the student-teacher consistency training scheme: the student produces standard outputs, while the teacher explicitly fuses its predictions with parts of the ground truth to give stronger supervision. We also exploit CMs' direct denoising for top-K multi-shot generation during training. Combining conditional generation with this enhanced consistency objective yields faster inference and improved prediction accuracy, establishing competitive new benchmarks on the large-scale Argoverse 2 dataset.
Alen Mrdovic, Qingze, Liu +6
Tony · Rutgers University, New Brunswick · The College of New Jersey