cs.CVJul 14, 2026

The GEST-Engine: From Event Graphs to Synthetic Video. A Full Technical Report

Authors: Nicolae CudlencoMihai MasalaMarius Leordeanu

Organizations: Institute of Mathematics of the Romanian Academy, Bucharest, Romania · Büchi Labortechnik AG, Flawil, Switzerland · National University of Science and Technology Politehnica Bucharest, Romania

Abstract

We present the GEST-Engine, a complete system that goes from natural-language text to fully-annotated multi-actor video. At its core is an explicit world model: rather than encoding state as a learned latent, the engine maintains a complete, inspectable representation of the world (which actors exist, where they are, what they are doing, which objects they hold, and how events relate in time and space), expressed as a formal Graph of Events in Space and Time (GEST) and realized deterministically inside the open world of a commercial game engine driven through an open-source multiplayer scripting framework. GESTs are produced either procedurally or by an agentic text-to-GEST system in which an LLM Director plans a story through tool calls validated by a programmatic state backend, so every generated specification is executable by construction. A GEST then enters a four-stage execution pipeline: graph parsing and validation, entity and action grounding, temporal orchestration (Allen-style constraints resolved by Floyd-Warshall transitive closure), and execution and capture. In a single simulation pass the engine emits frame-aligned RGB video, dense per-pixel depth, instance segmentation, per-actor skeletal pose, per-frame pairwise spatial-relation graphs, 2D bounding boxes, event-to-frame temporal mappings, and natural-language descriptions, all at zero marginal annotation cost. We further describe an in-game world editor, runtime capability extraction, a text-generation pipeline, and a production system that renders corpora at scale across parallel virtual machines. Because every frame traces back to a semantic specification, the engine guarantees object permanence, multi-actor coordination, and temporal consistency by construction, making its output valuable as training data, evaluation benchmarks, and diagnostic tools for video understanding.

Explore similar work

Apr 11, 2026cs.CV

Authoring for Living Worlds: Tool-Constrained LLM Agents for Executable Multi-Actor Scenarios

We use LLM agents to author executable specifications for a living world: formal Graphs of Events in Space and Time (GESTs) that a 3D game engine executes deterministically into multi-actor narrative videos, with per-frame spatial, temporal, and semantic ground truth as a byproduct of execution. This inverts the dominant paradigm of LLM agents driving neural video generators, which emit pixels with no semantic guarantees and no annotations. Authoring is the hard problem: the world's capability registry cannot be enumerated in a context window, validity of an action depends on accumulated world state, and a staged refinement pipeline driving GPT-5 through six validated stages produced zero executable specifications in 50 attempts. Our hierarchical Director / Scene Builder architecture instead operates through a constraint-enforcing tool layer, in which exploration tools paginate the registry and building tools validate every operation against simulator state, so every emitted specification is executable by construction. Driving a far smaller model (Claude Haiku 4.5), the system executes 20 of 25 attempts (80%) when seeded with a target narrative text. Because each seed text derives from a source graph, we can measure how faithfully the agent reconstructs specified intent: event-level F1 reaches 0.83 against a 0.55 matched-random floor, and sequential structure 0.77 against 0.43, with the residual gap dominated by information the text itself drops.
Nicolae Cudlenco, Mihai Masala, Marius Leordeanu
Apr 12, 2026cs.CV

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models

Game engines hold what video models struggle to learn: a complete, explicit world state behind every frame. We turn one into a data instrument. GEST-Engine, our production-grade open-source system, deterministically executes Graphs of Events in Space and Time (GESTs), whether procedurally generated or derived from text, into videos of synchronized multi-actor scenarios, recording ground truth as it renders: 3D entity and camera state, pairwise spatial relations, event-to-frame mappings, instance segmentation, and long descriptions, at zero marginal annotation cost. With it we release GTASA, a 938-video sample of what the system can generate at arbitrary scale, carrying, to our knowledge, the densest spatial-relation coverage of any video dataset: a complete entity-pair relation graph at every frame, ~84x denser than the state of the art, frame-for-frame. We validate GTASA both qualitatively, through human evaluation of physical validity and semantic alignment where frontier neural generators, given the same prompts, largely fail, and quantitatively, with GTASA pretraining improving VLM video captioning. Probing six frozen video encoders across 11 spatio-temporal tasks enabled by GTASA's exact 3D ground truth, a previously untestable inter-entity relational probe of frozen video features, reveals that who-is-near-whom barely rises above chance for all of them. We release the engine, the corpus, and the benchmark, making this gap a measurable, trainable target.
Nicolae Cudlenco, Mihai Masala, Marius Leordeanu
Mar 12, 2026cs.CV

Event-Driven Video Generation

Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact, an intended action is skipped, a placed object keeps drifting, or a support relation breaks. Our starting point is that standard frame-first denoising updates every latent region at every step, even when the prompt implies that only a local interaction should be active. We introduce Event-Driven Video Generation (EVD), a small DiT-compatible intervention that gives the sampler an explicit event signal. A lightweight head predicts token-level event activity; training losses tie that activity to latent state change; and event-gated sampling, with hysteresis and an early-step schedule, applies the update field mainly where an interaction is forming. On EVD-Bench, EVD improves human preference and VBench dynamics for state persistence, spatial accuracy, support relations, and contact stability, while keeping appearance quality comparable to the base model. The results suggest that a modest amount of event structure can correct several interaction failures that otherwise remain hidden behind good frame-level appearance.
Chika Maduabuchi, Jindong Wang