cs.CVJun 21, 2026

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories

Authors: Merve KocabasGege GaoBernhard SchölkopfAndreas Geiger

Organizations: University of Tübingen, Tübingen, Germany · Max Planck Institute for Intelligent Systems, Tübingen, Germany · ETH Zürich, Zürich, Switzerland · ELLIS Institute, Tübingen, Germany · Tübingen AI Center, Tübingen, Germany

Abstract

Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, while the intermediate generative dynamics remain hidden. Recent methods have begun to exploit generation order and process decomposition to improve sample quality, but still treat intermediate states as internal computation rather than objects for interaction. We propose Trajectory Forcing (TF), a trajectory-centric framework that makes the generation path explicit, semantic, and editable. TF organizes synthesis as a sequence of semantically structured stages, progressing from global layout to object-, part-, and detail-level representations. Each stage produces a decodable latent state that can be inspected, evaluated, and locally edited before the next stage begins. To instantiate this path, we derive coarse-to-fine teacher hierarchies by clustering pretrained visual representations such as DINOv2, and train a hierarchy-conditioned one-step flow-matching model at each level. We further introduce trajectory-aware metrics that measure structural consistency and local controllability beyond endpoint quality metrics such as FID. Experiments show that TF achieves competitive sample quality while exposing coherent intermediate states and supporting localized edits across semantic levels. By shifting the focus from final images to the generative path itself, TF opens a route toward controllable, trajectory-aware image synthesis.

Explore similar work

May 8, 2026cs.CV

Normalizing Trajectory Models

Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a few coarse transitions. Existing few-step methods address this through distillation, consistency training, or adversarial objectives, but sacrifice the likelihood framework in the process. We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training. Architecturally, NTM combines shallow invertible blocks within each step with a deep parallel predictor across the trajectory, forming an end-to-end network trainable from scratch or initializable from pretrained flow-matching models. Its exact trajectory likelihood further enables self-distillation: a lightweight denoiser trained on the model's own score produces high-quality samples in four steps. On text-to-image benchmarks, NTM matches or outperforms strong image generation baselines in just four sampling steps while uniquely retaining exact likelihood over the generative trajectory.
Jiatao Gu, Tianrong Chen, Ying Shen +3
May 11, 2026cs.LG

Follow the Mean: Reference-Guided Flow Matching

Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits a different control interface: adaptation through examples. For deterministic interpolants, the velocity field is solely governed by a conditional endpoint mean; shifting this mean shifts the flow itself. This yields a simple principle for controllable generation: steer a pretrained model by changing the reference set it follows. We instantiate this idea in two forms. Reference-Mean Guidance is training-free: it computes a closed-form endpoint-mean correction from a reference bank and applies it to a frozen FLUX.2-klein (4B) model, enabling control of color, identity, style, and structure while keeping the prompt, seed, and weights fixed. Semi-Parametric Guidance amortizes the same idea through an explicit mean anchor and learned residual refiner, matching unconditional DiT-B/4 quality on AFHQv2 while allowing the reference set to be swapped at inference time. These results point to a broader direction: generative models that adapt through data, not parameter updates.
Pedro M. P. Curvo, Maksim Zhdanov, Floor Eijkelboom +1
Mar 16, 2026cs.CV

WiT: Waypoint Diffusion Transformers for Alleviating Trajectory Conflict in Pixel-Space Image Generation

While recent Flow Matching models avoid the reconstruction bottlenecks of latent autoencoders by operating directly in pixel space, the raw pixel manifold provides little explicit semantic organization, making target-specific transport directions difficult for a shared finite-capacity vector field to distinguish. When different transport requirements are locally entangled in this way, their learning signals can interfere and hinder optimization, a phenomenon we refer to as trajectory conflict. To alleviate trajectory conflict while preserving direct generation in pixel space, we propose Waypoint Diffusion Transformers (WiT), which introduces explicit semantic routing into pixel-space generation. WiT structures pixel-space prediction through compact semantic waypoints projected from pre-trained vision representations, providing a semantically organized intermediate representation that complements direct pixel prediction. During ODE integration, a lightweight Waypoints Predictor dynamically infers the semantic waypoint from the current noisy state, and the predicted waypoint continuously conditions the primary Pixel Space Generator through our Just-Pixel AdaLN mechanism. This spatially varying semantic modulation reorganizes the finite-capacity learning problem and helps the model preserve target-specific transport directions while retaining generation entirely in pixel space. Experiments on ImageNet 2012 demonstrate consistent improvements over strong pixel-space baselines across model scales, with WiT-H/16 achieving an FID of 1.79. Code will be released at https://hainuo-wang.github.io/WiT.
Hainuo Wang, Mingjia Li, Xiaojie Guo