Many applications of physical AI run within finite or closed physical worlds with a limited set of physical laws governing object behavior. Examples include robots working in a warehouse and agents moving around in a video game. In order to better organize, filter, and curate data for physical AI applications, we propose a new approach centered on the unique settings and physical laws of individual datasets. We train LeCuration, a small world model intended to serve as a data curation tool for a separate, larger downstream model. To build this model, we choose LeWorldModel (LeWM)as our latent encoder and predictor, adding a diffusion transformer (DiT) decoder to add visuals to autoregressive gameplay rollout. We find that the embeddings of this model can be used as an anomaly detection signal and as a content-based clustering heuristic, and that auto-regressively predicting the game state with this model allows us to qualitatively check for action-state consistency. This paper presents a qualitative, proof-of-concept case study on CS:GO gameplay data; we do not yet report quantitative curation metrics or downstream training results, which we identify as the key next step.
Figures & tables
Figure 1: The full architecture of LeCuration. Frames are compressed to StableDiffusion latents using TAESD, which are then encoded as world state vectors by the latent encoder. The latent predictor takes 6 past world states and the next action vector to predict the next world state, which is supervised by encoding the next ground truth frame as a world state. The Diffusion Decoder generates TAESD latents for each gameplay frame individually, conditioned only on the predicted world state. These are decoded to RGB frames by the TAESD decoder.
Figure 2: Frames sampled from the dataset after HDBSCAN clustering. Each row is one cluster of world states in the latent space. Visual coherence within rows and clear differences between rows confirm that the learned embeddings recover semantically meaningful groupings of gameplay footage without any manual annotation. Additionally, the first row demonstrates how we can sift out anomalous clusters (in-scope views) without having to search for them specifically.
Figure 3: Animated embedding divergence during autoregressive rollout. Left : ground-truth CS:GO frame. Center : Each predicted world state decoded back to an RGB frame using the DiT. Right : Euclidean distance dt=∥e^t−et∥2 . Spikes identify frames where the predictor’s expectation diverges from reality, providing an unsupervised anomaly signal derived entirely from the world model’s own predictions.