Multiplayer world models must generate independently controlled views with consistent representations of both players and their shared environment. Most existing approaches coordinate multiple players through joint multi-view generation, whose cost grows with each additional player. We present WorldCast, a distributed multiplayer world model in which each player runs a local client comprising a video generator and a state model. Using recorded player positions and map geometry during training, the state model estimates the player's position from generated video and control inputs. Clients exchange player states and project them into camera-aligned player state fields that guide where and how other players are rendered. Shared scene state enables clients to reuse one another's generated observations to maintain consistent scene appearance across views. Experiments on Counter-Strike 2 demonstrate WorldCast's consistency, real-time performance, and distributed scalability. The camera-aligned player state field improves player rendering rates by over an order of magnitude over joint-generation methods, while shared scene state improves visual consistency over whole rounds. Each client runs in real time and exchanges only player and scene states, enabling scalable multiplayer generation without a centralized computational bottleneck. Image quality remains stable over hour-long rollouts.
Figures & tables
Figure 1: Three clients, each running on a separate machine, generate independently controlled views using shared world state. Only the initial frames are provided. Columns show successive moments, with controls annotated above each view; the bottom row shows a top-down layout fitted to the generated views. WorldCast supports real-time generation, cross-view consistency, scalability by adding clients, and distributed execution through state exchange.
Figure 2: (a) Coupled designs generate all players’ views in one model. (b) Each WorldCast player runs its own client, and clients exchange only the shared world state.
Figure 3: A client and the shared world state. Left: player states are projected into the camera-aligned player state field, and each generated block is added to the memory bank of the scene state, from which the next block retrieves one memory entry. Right: the causal video DiT of client i , where the player state field enters after the second DiT block, the memory frames join the recent and target frames, and every frame carries Plücker rays; after block n , the state model estimates the player’s position and the depth of each generated frame, which are published for block n+1 .
FVD ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Rendered ↑
Unmatched ↓
Frame concatenation
48.2
15.37
0.493
0.460
0.074
0.771
Solaris
49.0
15.66
0.501
0.449
0.056
0.739
WorldCast
39.7
16.17
0.507
0.418
0.817
0.039
Table 1: Bidirectional multiplayer generation on ten-second windows. All methods use the same backbone, training data, budget, and GT player positions.
Figure 4: Ten seconds of a held-out round, generated from the first frame and GT player positions (GT: recording with controls; Concat: frame concatenation).
First 30 s
After 30 s
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Trained without scene state
14.93
0.439
0.507
12.96
0.387
0.603
Scene state off at inference
14.37
0.424
0.545
12.93
0.391
0.615
Only its own earlier blocks
14.56
0.425
0.539
13.53
0.399
0.591
Least-covering memory
14.27
0.421
0.550
12.70
0.380
0.628
WorldCast
14.88
0.431
0.521
14.16
0.409
0.556
Table 3: Scene state over whole rounds. We use GT player positions to exclude the impact of imprecise player position in rendering performance.
Figure 5: State model on a recorded frame: (a) frame, (b) predicted depth, (c) predicted and GT trajectories and visibility over the map.
Figure 6: Serving P players on H200 GPUs: (a) FPS per player with decoding, (b) traffic received, (c) attention compute, (d) memory per GPU; Gamma-World and Solaris run on one GPU as released ((b): one player per GPU).
First 30 s
After 30 s
(a) Scene state
Off
On
Off
On
Mat. Pix. (K) ↑
72.0
92.1
31.1
44.3
Inliers ↑
46.2
51.5
18.8
27.7
RotErr ( ∘ ) ↓
31.8
29.0
59.1
55.6
Depth ↑
0.243
0.247
0.092
0.120
Table 4: (a) Closed-loop deployment over whole rounds, scene state off at inference and on: agreement between clients. (b) VBench ( Huang et al., 2024 ) over one-hour rollouts.
FVD ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Rendered ↑
Unmatched ↓
WorldCast
45.7
15.54
0.495
0.455
0.763
0.076
Trained without camera alignment
72.2
15.17
0.490
0.479
0.093
0.826
Trained without visibility
45.8
15.67
0.498
0.449
0.667
0.088
Trained without foreground weight
43.9
15.71
0.500
0.446
0.691
0.060
Trained with late injection
47.2
15.52
0.499
0.456
0.549
0.203
Trained with a coarse field
41.9
15.69
0.497
0.445
0.708
0.083
Table 5: Ablations of the player state field (stage 2, ten-second windows, 6,000 training steps each).
Figure 7: Closed-loop deployment: a whole round generated concurrently, each client on its own GPU and starting from its first recorded frame.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
0–10 s ↓
50–60 s ↓
Within two body widths at 60 s ↑
Blocked ↓
Recorded video
State model, Eq. ( 4 )
0.46
0.49
0.87
0.19
Extrapolation from the controls
1.86
7.63
0.02
2.38
Place head only
0.56
0.56
0.87
0.24
Motion head only
0.39
2.00
0.53
0.05
Generated video, WorldCast
Appendix
Table 6: Position error in body widths (32 u, 0.61 m) of the clients of the whole rounds of Table 3 , all alive at 60 s, on recorded video and on video generated by WorldCast with GT states: per-client medians over each ten-second interval, then medians over clients, the share of clients within two body widths at 60 s, and, as Blocked, the median error of the one-second displacement over the one-second windows in which the controls imply more than one body width of motion and the GT position moves less than a third of it. Bold marks the better of the state model and extrapolation; the last two rows of each part isolate the motion and place heads of the state model. Extrapolation does not read the video; its two rows differ only in the instants at which the error is evaluated.
Players P
2
4
8
10
16
FPS per player, including decoding ↑
WorldCast , one GPU per player
23.4
23.3
22.5
21.7
18.7
Gamma-World, one GPU
13.0
7.7
4.1
3.3
2.2
Solaris, one GPU
14.2
3.4
1.1
0.7
OOM
Memory per GPU, GiB
WorldCast , one GPU per player
27.1
27.0
27.0
27.1
27.1
Appendix
Table 7: Serving P players on H200 GPUs. WorldCast runs one client per player, each on its own GPU, with scene state, as deployed with CUDA graphs, which reproduce its output bit for bit. All P clients run concurrently and exchange messages and memory entries, through shared memory on one machine for P≤8 and through a shared file system across two machines for P=10 and 16 . A client’s frame rate counts its own time per block and excludes the time it waits for the other clients at the lockstep of our evaluations, since, without lockstep, a client reads the latest messages (Appendix B ); the table reports the median over clients and blocks. Gamma-World (2.2B parameters, 320×480 ; its distilled four-step causal student with a KV cache, the fastest configuration it releases) and Solaris (1.8B, 352×640 ) run as released, with random weights, all players on one GPU, which is their design. FPS per player includes decoding; memory per GPU is the peak of each framework’s allocator, for WorldCast the largest over clients. At P=16 , a WorldCast client receives 3.5 Mb/s on average: per block, a 12.0 kB message from each other client and, when it retrieves another client’s memory entry, that entry’s 387 kB of latents. Traffic is received per GPU; for Gamma-World and Solaris it is that of splitting the players one per GPU. Attention compute is the attention FLOPs needed to generate one second of video at 16 fps for all players on the GPU, computed analytically from the query–key products, 2d operations per pair, d being the model width; ∗ the configuration runs out of memory. OOM: the configuration exceeds the memory of one H200.
Stage
Initialization
Steps
Batch
Windows
1: bidirectional tuning, five-second windows
Wan2.2-TI2V-5B
34,000
256
8.70M
then ten-second windows
five-second model
6,000
256
1.54M
2: bidirectional training, player state field
stage 1
25,000
64
1.6M
then scene state
stage 2 at 20,000 steps
5,000
64
320k
3: autoregressive training
stage 2
5,000
64
320k
4: four-step distillation
stage 3
600
64
38.4k
Appendix
Table 8: Training runs of the reported models (Batch: global batch in windows); stages 3 and 4 are run once from each stage-2 model, with and without scene state. The ablations of Table 5 retrain the stage-2 model without scene state from stage 1 for 6,000 steps at a batch of 64 (384k windows). WorldCast denotes the stage-2 model without scene state at 25,000 steps in Table 2 , its 6,000-step checkpoint in Table 5 (the 384k column of Table 9 a), and the four-step model with scene state in Tables 3 , 4 , 6 , 7 , 15 , 16 and 17 – 19 .
Figure 8: Training stages, from a pretrained video model to four-step distillation.
(a) Training windows
64k
128k
256k
384k
512k
768k
1.0M
1.6M
Rendered ↑
0.617
0.667
0.738
0.763
0.759
0.774
0.784
0.817
Unmatched ↓
0.154
0.102
0.075
0.076
0.060
0.047
0.044
0.039
PSNR ↑
15.31
15.51
15.70
15.54
15.66
15.86
15.98
16.17
LPIPS ↓
0.469
0.456
0.448
0.455
0.444
0.436
0.428
0.418
FVD ↓
53.0
44.3
43.3
45.7
42.1
43.3
42.9
39.7
(b) Training windows
128k
256k
512k
768k
1.6M
Appendix
Table 9: Scaling with training data on the ten-second windows of Table 2 , with the same entity criterion. (a) The player state field. (b) The render rate of WorldCast , frame concatenation and Solaris at the budgets at which all three were trained and scored, each at its own checkpoint. Frame concatenation starts from a 5,000-step run (640k windows) without player states; its budgets count only windows with states.
First 5 s
Last 5 s
WorldCast
0.257
0.207
Frame concatenation
0.232
0.170
Solaris
0.233
0.179
Recorded video
0.317
0.309
Chance, depth from the most distant instant
0.117
0.116
WorldCast minus
Appendix
Table 10: Agreement between the depths of the two views of a ten-second window at the same instant, for the three models of Table 2 with GT states: share of pixels whose reprojected depths agree within the larger of 24 u and 5% , averaged over the windows; chance places a view’s depth from the most distant instant of its window at its current camera. Differences are paired over windows; their 95% bootstrap intervals over windows (10,000 draws) all exclude zero.
Δ PSNR
Δ SSIM
Δ LPIPS
Δ Rendered
Δ Unmatched
WorldCast minus frame concatenation
+0.80
+0.015
−0.042
+0.631
−0.571
WorldCast minus Solaris
+0.51
+0.006
−0.032
+0.636
−0.592
Predicted minus GT positions
−0.38
−0.007
+0.021
−0.058
+0.032
Appendix
Table 11: Paired differences on the ten-second windows of Tables 2 and 2 , whose 95% intervals from 10,000 bootstrap resamples of the windows, with the views of a window resampled together, all exclude zero. PSNR, SSIM and LPIPS report mean per-view differences, while Rendered and Unmatched report median per-view differences, which need not equal the difference between the medians of these tables. FVD compares two distributions of clips and therefore has no paired difference.
Δ Rendered
Δ Unmatched
Setting of the criterion
Concat.
Solaris
Pred.
Concat.
Solaris
Pred.
Match within half a body width
+0.648
+0.656
−0.095
−0.692
−0.709
+0.101
Match within two body widths
+0.553
+0.547
−0.038
−0.364
−0.409
+0.012
Detector confidence 0.25
+0.585
+0.584
−0.064
−0.603
−0.636
+0.029
Detector confidence 0.5
+0.556
+0.552
−0.090
−0.616
−0.624
+0.025
Boxes of at least 482 pixels
+0.638
+0.632
−0.053
−0.515
−0.521
+0.062
Appendix
Table 12: Table 11 under other settings of the entity criterion (Appendix D ), each changed in the generated and the recorded video alike: WorldCast minus frame concatenation (Concat.) and minus Solaris, and predicted minus GT positions (Pred.). Every difference keeps its sign, and its 95% bootstrap interval, computed as in Table 11 , excludes zero.
Subject
Background
Motion
Flicker
Dynamic
Imaging
Aesthetic
Whole rounds
0–10 s
0.732
0.888
0.953
0.926
1.000
70.8
0.480
10–20 s
0.792
0.917
0.965
0.947
1.000
71.1
0.491
20–30 s
0.822
0.929
0.966
0.948
0.969
71.5
0.499
30–40 s
0.824
0.926
0.966
0.948
0.948
71.8
0.497
40–50 s
0.827
0.929
0.964
0.946
0.958
72.1
0.497
Appendix
Table 13: Reference-free VBench scores of the ten-second windows of the whole rounds (GT states, four-step model trained without scene state) and of the one-hour rollouts of WorldCast with scene state, per twenty-minute interval. Dynamic degree is the fraction of clips classified as moving.
Damage
Subject
Background
Motion
Dynamic
Aesthetic
Imaging
Blur, σ=0.5 px
−0.1
+0.6
+0.3
0.0
−0.1
−5.7
Blur, σ=2 px
−1.1
+1.3
+1.1
0.0
−15.7
−48.1
JPEG, quality 30
−0.5
−2.1
−0.2
0.0
−7.7
−6.3
JPEG, quality 12
−1.6
−0.7
−0.3
0.0
−20.4
−19.6
Noise, σ=10
−0.8
−1.4
−0.8
0.0
−8.6
−14.9
Gamma 1.5
−0.1
−0.1
+0.1
0.0
−1.3
−7.4
Appendix
Table 14: VBench sensitivity: percent change under paired corruptions of random ten-second windows of the whole rounds, generated by the four-step model trained without scene state (Table 13 ).
Frames
Rendered by both
Disagreement (u) ↓
0–10 s
955
0.687
27.4
10–20 s
492
0.626
18.7
20–30 s
652
0.644
19.3
30–40 s
831
0.712
17.0
40–50 s
1,033
0.804
19.2
50–60 s
642
0.715
16.7
Appendix
Table 15: Agreement of two clients on the position of the player of the third client of their round, for WorldCast on the whole rounds with GT states. Frames denotes the number of frames in which that player is scored in both views, and Rendered by both the fraction of these frames in which both clients render it. Disagreement reports, over the frames rendered by both, the median closest distance between the two clients’ viewing rays to the persons matched to that player (one body width is 32 u).
Other players in view
1
2
3
4
Frames
44,395
15,274
3,498
631
Scored pairs
31,525
21,980
7,540
1,941
Rendered ↑
0.809
0.797
0.795
0.821
Appendix
Table 16: WorldCast on the whole rounds by the number of other players in view. Rendered is pooled over the scored pairs of each level of crowding, since one window mixes levels of crowding.
FVD ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Rendered ↑
Unmatched ↓
Trained without scene state
78.4
13.60
0.403
0.572
0.815
0.053
Scene state off at inference
90.5
13.40
0.402
0.593
0.820
0.049
Only its own earlier blocks
91.1
13.87
0.407
0.574
0.825
0.042
Least-covering memory
98.5
13.21
0.393
0.602
0.804
0.054
WorldCast
89.3
14.40
0.416
0.545
0.841
0.037
GT frames in the memory bank
75.0
14.99
0.433
0.513
0.853
0.030
Appendix
Table 17: Table 3 over whole rounds, against the recorded video. PSNR, SSIM and LPIPS: means over the ten-second windows; FVD: over their clips; Rendered and Unmatched: medians over views (Appendix D ).
WorldCast minus
Δ PSNR
Δ SSIM
Δ LPIPS
First 30 s
Trained without scene state
−0.05†
−0.0078†
+0.0138†
Scene state off at inference
+0.51
+0.0072†
−0.0248
Only its own earlier blocks
+0.32
+0.0059†
−0.0189
Least-covering memory
+0.61
+0.0097
−0.0293
GT frames in the memory bank
−0.29
−0.0113
+0.0163
Appendix
Table 18: WorldCast minus each other row of Table 3 , paired over the same ten-second windows: mean differences; † marks a 95% cluster-bootstrap interval over rounds (10,000 draws) that includes zero. Positive Δ PSNR and Δ SSIM and negative Δ LPIPS favor WorldCast . The last row, GT frames in the memory bank , whose memory entries hold GT frames that clients cannot have, is ahead (negative Δ PSNR and Δ SSIM, positive Δ LPIPS).
First 30 s
After 30 s
GT states
Trained without scene state
0.259
0.135
Scene state off at inference
0.229
0.133
Only its own earlier blocks
0.231
0.147
Least-covering memory
0.225
0.119
WorldCast
0.264
0.206
Appendix
Table 19: Agreement between the depths predicted for two clients’ generated frames at the same instant, on the whole rounds of Table 3 with GT states and, in the lower part, in the closed-loop deployment of Table 4 a with the cameras each client publishes: share of pixels whose reprojected depths agree within the larger of 24 u and 5% ; chance places a client’s depth from 10 s later at its current camera. Differences are paired over rounds; † marks a 95% cluster-bootstrap interval over rounds (10,000 draws) that includes zero. Own next latent frame applies the same measure to each client and its own next latent frame; scene state does not change it, so its gain between clients is not a steadier depth estimate.
Row minus WorldCast
Δ PSNR
Δ SSIM
Δ LPIPS
Δ Rendered
Δ Unmatched
Trained without camera alignment
−0.37
−0.005
+0.024
−0.545
+0.560
Trained without visibility
+0.13
+0.003
−0.006
−0.047
0.000†
Trained without foreground weight
+0.17
+0.006
−0.009
−0.043
0.000†
Trained with late injection
−0.02†
+0.004
+0.001†
−0.148
+0.051
Trained with a coarse field
+0.15
+0.003
−0.010
−0.034
0.000†
Field off at inference
−0.20
−0.003
+0.012
−0.690
+0.192
Appendix
Table 20: Each row of Table 5 minus WorldCast , paired over the views both rows score (PSNR, SSIM, LPIPS: mean difference; Rendered, Unmatched: median per-view difference, which need not equal the difference between the medians of Table 5 ); † marks a 95% interval from 10,000 bootstrap resamples of the windows, the two views of a window drawn together, that includes zero. Negative Δ PSNR, Δ SSIM and Δ Rendered and positive Δ LPIPS and Δ Unmatched mean that the row is behind WorldCast . FVD compares two distributions of clips and has no paired difference.
Figure 9: Nuke round 13, player 2; rows and overlays as in Fig. 4 .
Figure 10: Nuke round 13 of another match, player 1; rows and overlays as in Fig. 4 .
Figure 11: Nuke round 2, player 7; rows and overlays as in Fig. 4 .
Figure 12: Ancient round 22, player 5; rows and overlays as in Fig. 4 .
Figure 13: Ancient round 18, player 1; rows and overlays as in Fig. 4 .
Figure 14: Mirage round 16, one client, whole round with GT states. GT: recording; SceneOff: scene state off at inference; NoScene: trained without scene state. After 30 s, WorldCast keeps the scene of the recording and in the third column returns to the wall of its first frame; SceneOff and NoScene generate other parts of the map.
Figure 16: Ancient round 11, one player, ten-second window of Table 2 . GT: recording; Oracle: GT positions; Predicted: positions predicted in closed loop, every client estimating its own position from its generated video, with visibility from the depth head.
Figure 18: Mirage round 11, closed-loop deployment of Table 4 a: all five players of one team run clients with predicted states; three are shown. Off/On: scene state off/on; a player’s Off and On rows differ only in scene state. With scene state, the three views stay around the A-site ramp and agree on its landmarks, and Players 2 and 3 draw the teammates ahead of them (supplementary videos).
Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsistencies, and poor scalability. We propose MASS (Multiplayer world models with Authoritative Shared State) to resolve this limitation. Inspired by multiplayer game architectures, MASS disentangles world dynamics and view rendering. A learned Logic Engine advances a global, authoritative typed state from joint actions without any hand-written transition function, acting as the sole recurrent memory and synchronization reference. From this shared state, a learned Rendering Engine generates independent and consistent views for any requested camera on demand. This explicit disentangling allows MASS to achieve superior state accuracy and lower cross-view inconsistency compared to state-of-the-art multi-view baselines on a matched multiplayer Snake benchmark. It advances predicted worlds with 1,024 concurrent players for 10,000 recurrent steps. Our results show that explicit, authoritative state modeling provides a practical foundation for scalable and consistent multi-agent world simulation.
Ziqi Cai, Siqi Yang, Yimu Wang +6
1Alaya Lab 2Peking University 3Institute of Science Tokyo
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly available bots, our 5-billion-parameter latent diffusion model generates four-player matches in real time, producing 20 frames per second on a single Nvidia B200 GPU. Although trained only on short clips, its rollouts stay stable far beyond the training horizon: distributional quality holds steady out to five minutes, the longest horizon we measure, and in practice we observe rollouts continuing for hours with no sign of collapse. We systematically investigate the central design choices: the video codec, the generative objective, and the multiplayer conditioning scheme. In addition, we characterize how behavior changes with model and data scale, including the capabilities that emerge and the failure modes that persist. We further develop targeted evaluations that probe the model's physical understanding rather than visual appearance alone. To support continued research on multiplayer world models, we release our dataset, our full training and inference codebase, and a live demo.
Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary +24
1General Intuition · 2Kyutai · ‡École nationale des ponts et chaussées +1
Multiplayer world models must ensure that independently controlled views remain consistent with one shared and persistent world. We introduce MultiWorldBench, a diagnostic Minecraft benchmark containing 495 case configurations across seven task suites and ten capabilities, including independent control, cross-view motion, shared-state synchronization, persistence, structural reasoning, concurrent interaction, and delayed revisit. We evaluate Solaris, Gamma-World, and MineWorld, using Engine GT as a reference. Gamma-World achieves the highest ten-capability average among the generated systems at 21.39, followed by Solaris at 20.88 and MineWorld at 1.89, while Engine GT reaches 91.69. Gamma-World performs better on several control, shared-state, and revisit capabilities, whereas Solaris leads in cross-view motion and race-condition consistency. Nevertheless, all generated systems score at most 8.00 on state persistence and 1.33 on structural consistency, and none succeeds in spatial reasoning or building-identity preservation. Human preferences produce the same overall ranking and show strong alignment with the automatic evaluation, with a mean dimension-level Spearman correlation of 0.96. These results show that plausible individual views do not yet constitute a coherent multiplayer world.