cs.CVOct 8, 2026

Dino Forcing Flow Models: Do not denoise what you can predict

Authors: Arijit Ghosh, Lucas Degeorge, Paul Couairon, Alexei A Efros, Vicky Kalogeiton, David Picard

Organizations: ENPC, IP Paris · Ecole Polytechnique, IP Paris · AMIAD · UC Berkeley

Abstract

Co-denoising pretrained representations such as DINO can substantially improve the training speed and quality of flow matching models, but it introduces a second denoising trajectory and requires carefully designed schedules. We propose a simpler alternative: predict the pretrained representation directly, then condition the model on its own prediction. This removes the need for a second ODE and any representation-specific denoising schedules, while retaining the benefits of representation guidance. Our approach converges substantially faster and achieves better generation quality as measured by FID score. On ImageNet, it outperforms the state of the art in latent space at 2x fewer epochs than prior methods; in pixel space, it improves FID over comparable prior methods by more than 20%. These results support a simple principle: do not denoise what you can predict. Our code is openly available at https://github.com/arijit-hub/dino_forcing.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RiT: Vanilla Diffusion Transformers Suffice in Representation Space

    May 21, 2026Le Zhang, Ning Mang, Aishwarya AgrawalFlow MatchingVisual Representation Learning

  2. Rethinking Pixel Mean Flows via Interval Denoiser

    Aug 5, 2026Alexander Zaytsev, Dmitry Baranchuk, Alexander Korotin +1Flow MatchingImage Generation

  3. Training Flow Matching: The Role of Weighting and Parameterization

    Mar 6, 2026Anne Gagneux, Ségolène Martin, Rémi Gribonval +1Flow MatchingAdaptive Loss Weighting