Benchmarks for measuring the quality of action-conditioned world models are still evolving and shifting away from visual similarity-based metrics to action-semantic and physically-grounded metrics. However, for domain and task-agnostic action-conditioned world model training, existing benchmarks provide a limited signal. By training and evaluating diffusion world models on CounterStrike gameplay data, we confirm that qualitative playability does not correspond with metrics such as FVD, LPIPS, and JEDi. We term this the Virtual2Real gap. We posit that, in lieu of reliable benchmarks, curating raw gameplay data and measuring a variety of diagnostic properties provides a more robust signal to bridge the gap, before the training even begins. We present several curation strategies and a general-purpose kit for physical AI data curation called Kuration SDK, which is being open-sourced with this paper. The SDK was instrumental in uncovering the root cause of the virtual2real gap in a specific case: why two world models trained on identical gameplay map, action and state distribution, behaved very differently when played in spite of having very similar LPIPS and FVD scores. Thus, Kuration SDK has the potential to uncover the root causes of Virtual2Real gap in specific datasets and accelerate development of sample-efficient training datasets.
Figures & tables
Run
Epoch
LPIPS ↓
FVD ↓
DIAMOND (paper, reported)
600
0.5408±0.1829
1,643,391
DIAMOND weights (our eval)
600
0.5414±0.1836
1,655,984
Ours (best FVD)
300
0.5789±0.1839
1,571,046
Ours (best LPIPS)
420
0.5634±0.1916
1,801,407
Table 1: Training parity with DIAMOND on the published CS:GO dataset. All evaluations use 500 test episodes and 1-step samplers.
Figure 1: LPIPS, FVD, and JEDi for baseline_v1 across training, evaluated on the CS2 test set, plotted on three shared axes. Dashed lines show the DIAMOND model on the same test set (no JEDi reference exists on this test set). LPIPS improves with noise; FVD is competitive only at early epochs and then oscillates; JEDi improves at every evaluated checkpoint but playability is still worse.
Figure 2: Action-state consistency for single key-tap movement controls, DIAMOND’s public CS:GO corpus vs. the partner CS2 corpus. (a) CS:GO responses are consistently larger in magnitude. (b) CS:GO responses take substantially longer to decay to zero after release, evidence of sustained momentum absent from the CS2 corpus.
Strategy
Epochs
Init.
LPIPS ↓
FVD ↓
JEDi ↓
Playability
DIAMOND (our eval)
600
scratch
0.6049
2,974,382
—
better
baseline_v1 (uncurated)
600
scratch
0.4785
2,898,653†
2.39
reference
Coverage-based
120
scratch
0.5404
3,043,786
—
no change
Rare-trajectory
60
epoch 300
0.5043
—
3.13
no change
Collision-aware
55
epoch 300
—
—
2.76
no change
Action-state-consistency
—
—
—
—
—
improved
Table 2: DIAMOND and curation strategies against the uncurated baseline_v1 control, all evaluated on the CS2 test set (1,090 episodes). Coverage-based training is from scratch; rare-trajectory and collision-aware are fine-tuned from baseline_v1 ’s epoch-300 checkpoint (see Appendix A for full per-epoch curves; Section 7.1 for coverage-based’s epoch-105 best checkpoint). Action-state-consistency’s training details are not yet finalized (Section 7.2 ). Playability is a qualitative judgment from interactive rollouts, relative to baseline_v1 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: LPIPS (left axis) and FVD (right axis) at each 30-epoch checkpoint of our 600-epoch DIAMOND reproduction on CS:GO, plotted on a shared axis. Dashed lines mark DIAMOND’s reported values. FVD crosses below the paper’s reported value at epoch 180 and reaches its best value at epoch 300; after that, both metrics plateau and oscillate without a consistent trend even though training loss continues to decrease monotonically.
Epoch
LPIPS ↓
FVD ↓
JEDi ↓
30
0.6910±0.1518
3,145,978
5.5917
60
0.6862±0.1634
4,358,746
4.9600
90
0.6557±0.1716
3,631,473
3.5018
120
0.6517±0.1749
2,349,638
2.9815
150
0.6157±0.1778
2,374,627
2.9877
180
0.6204±0.1823
1,627,126
3.5083
Appendix
Table 3: LPIPS, FVD, and JEDi for the DIAMOND reproduction at each 30-epoch checkpoint, DIAMOND standard CS:GO test set.
Epoch
LPIPS ↓
FVD ↓
JEDi ↓
30
0.5734±0.1916
4,313,603
6.3680
60
0.5515±0.2077
2,898,653
5.5596
90
0.5579±0.2081
4,373,784
4.0920
120
0.5203±0.2112
3,411,944
4.4658
150
0.5277±0.2075
3,714,381
3.5952
180
0.5040±0.2041
3,113,235
4.3480
Appendix
Table 4: LPIPS, FVD, and JEDi for baseline_v1 at each evaluated checkpoint on the CS2 test set.