cs.CVSep 30, 2026

CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models

Authors: Shu Yu, Chaochao Lu

Organizations: Shanghai Artificial Intelligence Laboratory, Shanghai, China · Shanghai Innovation Institute, Shanghai, China · Fudan University, Shanghai, China

Abstract

Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement. To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07x that of Flow-GRPO, while overall generation quality also improves.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models

    May 11, 2026Andreas Bergmeister, Stefanie Jegelka, Nikolas Nüsken +2Generative Flow NetworksDiffusion Models

  2. Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models

    Sep 29, 2025Shuchen Xue, Chongjian Ge, Shilong Zhang +2Diffusion ModelsDiffusion Policies

  3. DRM: Diffusion-based Reward Model With Step-wise Guidance

    May 25, 2026Jaxon Zhang, Binxin Yang, Hubery Yin +2Diffusion AlignmentDiffusion Models