cs.LGMay 14, 2026

ff-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data

Authors: Jake FawkesJason Hartford

Organizations: Department of Statistics, University College London, UK · Valence Labs, London, UK · Recursion, London, UK.

Abstract

In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities is an effective, low variance, surrogate loss for training generative models. This loss has the property that when evaluated \emph{on-policy} its gradients correspond to those of the KL divergence, while \emph{off-policy} it remains a valid loss with the same global minimizer. In this work, we demonstrate that this construction can be extended to the whole family of ff-divergences, leading to a family of losses whose on-policy gradients are that of the corresponding ff-divergence, but retain the same global minimizer off-policy. Specifically, we show that the on-policy gradients lead to a one to one correspondence between translation invariant loss functions on the target and model log probabilities, and ff-divergences. This equivalence allows us to design new surrogate loss functions for tuning a wide class of generative models that inherit the properties of the corresponding ff-divergence, such as being more mode covering, whilst being applicable to off-policy data. We apply our losses on a range of tasks, including classic synthetic examples, SynFlowNets for molecule discovery, and asynchronous large language model (LLM) tuning, demonstrating that our models retain their predicted properties on- and off-policy in a wide class of generative models.

Explore similar work

CardsList
  1. GFlowNets and variational inference

    Oct 2, 2022Esmeralda S. Whitammer, Salem Lahlou, Tristan Deleu +5Variational InferenceGenerative Flow Networks

  2. Stable GFlowNets with Probabilistic Guarantees

    May 3, 2026Zengxiang Lei, Ananth Shreekumar, Jonathan Rosenthal +6Generative Flow NetworksDistributional Learning