Crowd simulation plays a central role in robot navigation, autonomous driving, and urban planning. For these applications, realistic simulation requires crowds to adapt their behavior to environmental changes and user objectives. However, existing methods that rely on predefined control settings have limited flexibility in accommodating new user-specified objectives. To address this limitation, we propose Ctrl-CWM, a multi-agent Controllable Crowd World Model that integrates crowd generation and run-time control. Our key idea is to adapt the world-model principle of planning using imagined futures to crowd simulation. To this end, Ctrl-CWM consists of an encoder that learns a representation of human motion dynamics, an actor that proposes pedestrian displacements, a critic that evaluates imagined crowd trajectories, and a planner that selects actions. We first learn human motion dynamics through trajectory prediction on real-world pedestrian videos and then freeze the encoder to preserve them. Using this representation, the actor generates imagined crowd trajectories through repeated state updates, and the planner combines the critic's scores with user costs to select actions. Repeated planning advances the simulated crowd, while additional user costs introduce new control objectives without retraining. We extensively evaluate crowd generation under varied agent arrival conditions and run-time control across avoidance and attraction scenarios. Ctrl-CWM outperforms the state-of-the-art method on most crowd realism and collision metrics, and adapts crowd behaviors to user-specified objectives introduced during simulation. The project page is available at https://jungyu0413.github.io/Ctrl-CWM
Figures & tables
Figure 2: Overview of Ctrl-CWM. Trajectory prediction trains the state encoder hθ on real-world pedestrian data, which is then frozen. The actor πψ and critic Vϕ are learned over imagined rollouts on this representation, and CEM adds a user cost to the critic score for run-time control.
Arrivals from the RS emitter
Dataset/Fold
Model
Scene-Level Realism
Agent-Level Accuracy
Dens. ↓
Freq. ↓
Cov. ↓
Pop. ↓
Kinem. ↓
DTW ↓
Div. ↑
Col.(%) ↓
ETH
ORCA
0.019
0.016
0.016
0.205
0.839
2.848
0.092
0.004
CrowdES
0.044
0.025
0.025
0.461
1.138
2.608
0.240
2.660
Ours
0.034
0.016
0.016
0.398
0.358
1.511
0.171
0.265
HOTEL
ORCA
0.044
0.017
0.017
0.396
0.623
1.094
0.179
0.015
Table 1: Crowd behavior generation under shared arrivals. Results use identical replayed random-surface (left) and diffusion (right) arrivals across methods. AVG averages five ETH–UCY folds, SDD, and GCS. Lower is better except Div.; Col. is in %. Bold / underline : best/second-best.
Table 3
Figure 4: Run-time control. The left and right pairs show uncontrolled (red) and controlled (green) rollouts, respectively, at t1 and t2 . (a) An avoidance objective redirects pedestrians away from the hazardous zone. (b) An attraction objective guides them toward the fountain.
Variant
Scene-level realism
Agent-level accuracy
Control
Dens. ↓
Freq. ↓
Cov. ↓
Pop. ↓
Kinem. ↓
DTW ↓
Div. ↑
Col.(%) ↓
Comp. ↑
Random encoder
0.101
0.067
0.067
0.304
1.205
1.031
0.157
1.101
N/A
Actor only
0.077
0.047
0.047
0.271
0.432
0.970
0.280
0.931
N/A
User-cost only
0.089
0.075
0.075
0.294
0.812
1.178
0.078
2.398
0.858
Ctrl-CWM (full)
0.078
0.047
0.047
0.254
0.465
0.925
0.285
0.756
0.824
w/o soft-DTW
0.082
0.051
0.051
0.275
0.846
0.989
0.182
0.838
N/A
Table 3: Component ablations under diffusion arrivals. Metrics are averaged over five ETH–UCY folds. User-cost only retains the actor, imagined rollouts, and CEM procedure but ranks candidates using only the user cost. Comp. reports avoidance compliance Cavoid under the same protocol. N/A denotes generation-only evaluation. Bold / underline : best/second-best.
Figure 5: Ablation of planning parameters. Lines show compliance and trajectory metrics over the post-activation control window. Gray bars report call latency in milliseconds on the right axes. Dotted lines mark the default settings.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Operation
Input
Output
Inputs and scene encoder
1-1
Similarity transform
Scene, trajectories
Grid, G=192
1-2
Walkability and semantic channels
Transformed scene
9×192×192
1-3
Focal-agent Gaussian, σ=5
Agent position
1×192×192
1-4
Crowd Gaussian accumulation
Crowd positions
1×192×192
1-5
Channel concatenation
1-2, 1-3, 1-4
11×192×192
Appendix
Table 4: Encoder and prediction-head interfaces. Spatial dimensions are per focal agent, with batch axes omitted. A is the number of co-present agents and Th is the velocity-history length. Widths use w=0.25 . GN(8) denotes GroupNorm with eight groups.
Module
Layers or input
Output
Actor input
[ℓit∥cit∥uit∥vit/G]
108 features
Actor hidden
Linear 108→128→128 , ReLU
128 features
Actor output
Linear 128→2
Displacement
Critic input
Six geometric features at each imagined step
H×6
Critic recurrent
GRU 6→128 , final hidden state
128 features
Critic output
MLP 128→128→1
Rollout score
Appendix
Table 5: Behavior-learning networks. The actor receives frozen encoder features in the full model. The critic uses H imagined steps, while its training targets use L steps. The discriminator uses L displacement vectors.
Setting
Value
Prediction pretraining
Optimizer
AdamW
Learning rate / weight decay
3×10−4 / 10−5
Learning-rate schedule / gradient clipping
Cosine / 1.0
Epochs / scenes per optimizer update
80 / 3
Batch size / accumulation steps
1 / 3
Appendix
Table 6: Training and planning settings. The three calibrated weights are not fixed to the calibration ratio. User-cost weights are reported with each control protocol.
Evaluation
Setting and aggregation
Repeats and window
Generation, Table 1
Two replayed arrival protocols. Mean of five ETH–UCY folds, SDD, and GCS.
Full episodes.
Avoidance, Table 2
Disc, rectangle, multiple zones. λuser=10 . Same seven-entry mean.
10 trials per scene. Final two thirds.
Components, Table 3
Diffusion arrivals over five folds. Generation ablations use full episodes; the user-cost-only comparison uses the same avoidance-control protocol as the full model.
Full episodes for generation; post-activation window for control.
Planning, Fig. 5
Five folds. K , I , and H varied individually.
Post-activation window.
Additional components, Appendix C
Five folds. Disc avoidance, λuser=10 .
20 trials per scene. Final two thirds.
Weight/population sweeps, Appendix D
Five folds. Weights and emitter multipliers varied as reported.
One trial per configuration. Post-activation window.
Appendix
Table 7: Evaluation protocols. Repeat counts are shown where reported. The generation and control columns of Table 3 use separate evaluations.
Full model
w/o interaction
Fold
Col. (%) ↓
Comp. ↑
Col. (%) ↓
Comp. ↑
ETH
0.458
0.932
20.32
0.950
HOTEL
1.522
0.992
1.292
0.497
UNIV
0.766
0.882
1.133
0.468
ZARA1
3.508
0.839
0.868
0.466
ZARA2
2.033
0.754
1.062
0.510
Appendix
Table 8: Interaction-encoder ablation. Disc avoidance with λuser=10 and 20 trials per scene. AVG is calculated from the displayed fold values.
Fold
Full model
w/o dynamics head
ETH
0.932
0.880
HOTEL
0.992
0.215
UNIV
0.882
0.245
ZARA1
0.839
0.880
ZARA2
0.754
0.699
AVG
0.880
0.584
Appendix
Table 9: Auxiliary dynamics-head ablation. Avoidance compliance under the same disc-control protocol. AVG is calculated from the displayed fold values.
Variant
Kinem. ↓
DTW ↓
Col. (%) ↓
Comp. ↑
Pointwise critic
0.433
0.994
2.454
0.870
Full model (listwise)
0.404
0.985
1.657
0.880
Appendix
Table 10: Critic-objective comparison. Five-fold means under the 20-trial disc-control protocol.
Avoidance compliance ↑
Fold
λ=5
λ=10
λ=25
λ=50
λ=100
ETH
0.657
0.949
0.987
0.955
0.997
HOTEL
0.661
0.823
0.990
0.964
0.992
UNIV
0.694
0.931
1.000
1.000
1.000
ZARA1
0.422
0.759
0.963
0.973
0.938
ZARA2
0.600
0.855
0.986
0.949
0.967
Appendix
Table 11: Avoidance weight sweep. Five folds, one trial per configuration, post-activation evaluation. Kinematic changes are controlled minus uncontrolled. λ is λuser .
Fold
λuser=1.5
λuser=3.0
λuser=5.0
ETH
0.480
0.674
0.977
HOTEL
0.406
0.580
0.907
UNIV
0.244
0.670
0.960
ZARA1
0.390
0.572
0.930
ZARA2
0.386
0.727
0.983
AVG
0.381
0.645
0.951
Appendix
Table 12: Attraction weight sweep. Avoidance and attraction weights have different cost scales. One trial is used per configuration. AVG is calculated from the displayed fold values.
λuser=3.0
λuser=5.0
Fold
Δ Kinem.
Δ Col. (pp)
Δ Kinem.
Δ Col. (pp)
ETH
+13.92
+11.2
+13.78
+37.9
HOTEL
+5.41
−0.4
+10.66
+22.2
UNIV
+3.14
+3.0
+3.45
+32.8
ZARA1
+4.09
−0.5
+4.98
+15.6
ZARA2
+5.63
+0.7
+5.93
+24.1
Appendix
Table 13: Motion changes under attraction. Post-activation controlled minus uncontrolled values from the supplied sweeps. Collision changes are percentage points (pp), not relative percentages.
λuser=10
×0.5
×1
×2
Fold
Comp. ↑
Δ Col. (pp)
Comp. ↑
Δ Col. (pp)
Comp. ↑
Δ Col. (pp)
ETH
0.928
+9.5
0.949
−0.0
0.836
−0.1
HOTEL
0.843
+0.0
0.823
+0.0
0.632
+0.0
UNIV
1.000
−0.0
0.931
−0.0
0.732
+0.0
ZARA1
0.651
−0.2
0.759
−0.2
0.744
+0.5
Appendix
Table 14: Emitter-population sensitivity. Compliance and post-activation collision-rate changes at each emission multiplier. One trial is used per configuration. All signs, including rounded zero changes, follow the source.
Model
Year
ETH
HOTEL
UNIV
ZARA1
ZARA2
AVG
SDD
GCS
AgentFormer Yuan et al. (2021)
2021
0.46/0.80
0.14/0.22
0.25/0.45
0.18/0.30
0.14/0.24
0.23/0.40
8.7/14.9
10.2/16.9
MID Gu et al. (2022)
2022
0.57/0.93
0.21/0.33
0.29/0.55
0.28/0.50
0.20/0.37
0.31/0.54
7.6/14.3
10.7/18.2
EqMotion Xu et al. (2023)
2023
0.40/0.61
0.12 /0.18
0.23/0.43
0.18/0.32
0.13 /0.23
0.21 /0.35
7.9/11.9
7.6/13.1
MART Lee et al. (2024)
2024
0.35 /0.47
0.14/0.22
0.25/0.45
0.17/ 0.29
0.13 / 0.22
0.21 /0.33
7.4 /11.8
10.6/14.1
LMTraj Bae et al. (2024)
2024
0.41/0.51
0.12 / 0.16
0.22 / 0.34
0.20/0.32
0.17/0.27
0.22/0.32
7.8/ 10.1
7.1 / 9.6
MoFlow Fu et al. (2025)
2025
0.40/0.57
0.11 /0.17
0.23/0.39
0.15 / 0.26
0.12 / 0.22
0.20 /0.32
7.5/12.0
9.1/11.6
Appendix
Table 15: Trajectory prediction on ETH–UCY, SDD, and GCS. Best-of-20 ADE/FDE is reported in meters for ETH–UCY and pixels for SDD and GCS. AVG is the mean over the five ETH–UCY folds. Bold / underline : best/second-best. N/A denotes unreported results.
Figure 6: Run-time control in additional reconstructed scenes. Columns show the initial scene, motion without control (red), and motion with control (green). Each objective is illustrated in three scenes. (a) Avoidance redirects pedestrians away from a hazardous region. (b) Attraction guides them toward a specified target.
Data
Method
Dens. ↓
Freq. ↓
Cov. ↓
Pop. ↓
Kin. ↓
DTW ↓
Div. ↑
Col. ↓
ETH
ORCA
0.019
0.016
0.016
0.205
0.839
2.848
0.092
0.004
CrowdES
0.044
0.025
0.025
0.461
1.138
2.608
0.240
2.660
Ctrl-CWM
0.034
0.016
0.016
0.398
0.358
1.511
0.171
0.265
HOTEL
ORCA
0.044
0.017
0.017
0.396
0.623
1.094
0.179
0.015
CrowdES
0.015
0.016
0.016
0.130
0.524
0.907
0.184
2.471
Ctrl-CWM
0.012
0.012
0.012
0.098
0.419
0.746
0.224
0.828
Appendix
Table 16: Random-surface-arrival generation results. Complete values from Table 1 . Col. is a percentage. Bold / underline identify the best/second-best distinct displayed values within each benchmark entry.
Data
Method
Dens. ↓
Freq. ↓
Cov. ↓
Pop. ↓
Kin. ↓
DTW ↓
Div. ↑
Col. ↓
ETH
ORCA
0.026
0.013
0.013
0.265
0.867
3.794
0.096
0.006
CrowdES
0.020
0.020
0.020
0.208
0.377
1.649
0.203
0.697
Ctrl-CWM
0.017
0.006
0.006
0.193
0.214
1.458
0.181
0.620
HOTEL
ORCA
0.022
0.010
0.010
0.199
1.052
1.588
0.112
0.013
CrowdES
0.013
0.009
0.009
0.117
0.336
0.643
0.242
1.197
Ctrl-CWM
0.011
0.009
0.009
0.087
0.305
0.548
0.259
1.046
Appendix
Table 17: Diffusion-arrival generation results. Complete values from Table 1 . Col. is a percentage. Bold/underline identify the best/second-best distinct displayed values within each benchmark entry.
Disc
Data
ORCA + obstacle
CrowdES + map
Ctrl-CWM
ETH
1.000
0.599
0.837
HOTEL
-1.154
0.546
0.988
UNIV
0.530
0.727
0.831
ZARA1
0.428
0.469
0.752
ZARA2
1.000
0.598
0.714
Appendix
Table 18: Full avoidance-control results. Each zone block reports Cavoid . Baseline obstacles are present from initialization, while Ctrl-CWM activates the objective after a shared first third. AVG is the unweighted mean of the seven entries.
Safe and efficient robot navigation in crowds requires anticipating pedestrian motion despite uncertain and potentially shifting prediction errors. Existing reactive methods can produce oscillatory behavior, while predictive planners often treat forecasts as exact or rely on restrictive error models. Incorporating conservative uncertainty sets as hard constraints can also render model predictive control (MPC) infeasible. We propose \textit{CoCoNav}, a crowd-navigation framework that combines online conformal calibration with runtime-certified planning. A horizon-specific conformal proportional--integral controller adapts trajectory-error bounds to regulate long-run empirical coverage, enabling the framework to respond to changing prediction errors. A \textit{relax-then-verify} planner preserves solver feasibility by generating nominal trajectories with soft-constrained MPC and separately certifying them, together with contingency maneuvers, against the calibrated bounds before execution. Simulations and quadruped experiments show that CoCoNav achieves a favorable balance among collision avoidance, task success, and navigation efficiency relative to the evaluated baselines.
Cheng Guo, Mingzhe Ni, Zheng Liang +5
Department of Computer Science, The University of Manchester, Manchester, UK. · Human-Robot Interfaces and Interaction Laboratory, Italian Institute of Technology, Genoa, Italy. · Genisom AI, Shanghai, China. +3
Controllable traffic simulation is critical for testing autonomous driving systems, yet existing approaches often require retraining large generative models with extensive annotated data. We introduce a lightweight control adaptation framework that enables multi-modal controllability (sketch, latent behavior codes, and text) for pretrained state-of-the-art diffusion and autoregressive traffic models. By modulating intermediate features through identity-initialized FiLM layers, our method efficiently adds new control modalities while preserving the base model's generative prior. Evaluated on Waymo Open Sim Agents Challenge, our approach demonstrates strong controllability with less than 1% of the paired control data. Through context-aware condition transfer, our framework enables counterfactual scenario generation and long-tail synthesis while maintaining stable closed-loop driving realism and safety. Our framework unlocks new possibilities for controllable traffic simulation, enabling targeted scenario generation through lightweight adaptation of pretrained generative models. Project page: https://ecosim-web.github.io/
Yu-Hsiang Chen, Wei-Jer Chang, Yi-Ting Chen +1
National Yang Ming Chiao Tung University, Taiwan · University of California, Berkeley, USA
Autonomous mobile service robots are often required to complete tours that require navigating through a set of locations in an environment. Example domains include guiding people through a shopping mall, delivering packages in a fulfilment centre, or giving guided tours in a museum. However, in crowded environments, the presence of people may negatively impact robot performance. For example, humans will activate robot collision avoidance manoeuvres that slow the robot down. Crowds move stochastically and vary throughout the day. In this paper we present a probabilistic tour planner for crowded environments which explicitly reasons over human congestion. We learn circular linear flow field (CLiFF) maps which predict human trajectories given an initial observation. We then use these predictions to build and solve a Markov decision process online which efficiently routes the robot through the environment. Our approach is scalable enough to re-plan as new people are observed. We evaluate our approach on a real-world crowd dataset in a shopping mall.
Stefano Bernagozzi, Charlie Street, Masoumeh Mansouri +1
1Istituto Italiano di Tecnologia – Genova, Italia · 2Universit`a di Genova – Genova, Italia · University of Birmingham – Birmingham, United Kingdom