Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.
Figures & tables
Fig. 1: Overview of Uni-TMPO for T2I generation and VLA manipulation. Progress-conditioned coarse-to-fine sampling constructs T2I trajectory groups for each prompt, while feedback-conditioned action-chunk sampling constructs VLA trajectory groups from a shared initialization. Unlike scalar reward maximization, Uni-TMPO matches reward-induced and policy-induced allocations within each group, preserving diverse image modes and action strategies. The right panels summarize T2I performance and VLA OOD generalization.
Fig. 2: Five-mode distribution matching with a three-layer MLP. The left panels show the five-peak reward landscape and shared initialization. The remaining panels show intermediate and final samples and final mode mass under Uni-TMPO (top) and Reward Maximization (bottom). Uni-TMPO matches the target and preserves all five modes, whereas Reward Maximization collapses to the highest-reward mode.
Method
Time (s) ↓
GenEval ↑
OCR ↑
PickScore ↑
HPS ↑
ImgRwd ↑
LGMD ↑
Cos.Div. ↑
Compositional Image Generation (GenEval)
FLUX.1-dev
–
0.647
–
22.301
0.301
1.099
−0.031
0.211
DAG-DB
187.5
0.889
–
21.998
0.291
1.071
0.097
0.237
DGFS-SubTB
178.6
0.917
–
22.210
0.298
1.107
0.113
0.241
Flow-GRPO
160.8
0.946
–
22.113
0.289
1.074
−0.089
0.198
TreeGRPO
126.2
0.936
–
21.524
0.281
1.083
−0.281
0.184
TABLE I: Comparison of FLUX.1-dev T2I post-training across compositional image generation, visual text rendering, and human preference alignment. The task-specific training reward is shown in parentheses after each task. Best and second-best results are bolded and underlined.
Fig. 3: T2I reward–diversity–efficiency trade-off. Metrics are normalized across the post-training methods within each task. Larger values indicate better results on every axis; the Time axis is reversed because shorter iteration time is better. Uni-TMPO provides the strongest overall trade-off across compositional generation, visual text rendering, and human preference alignment.
Fig. 4: T2I diversity under matched prompts. Each panel compares three stochastic samples per method under the same prompt. Compared with Flow-GRPO, Uni-TMPO produces broader variation in object appearance, aesthetic style, and spatial layout. Table I reports quantitative results.
Backbone
Method
LIBERO
MetaWorld-MT50
CALVIN-D
Spatial
Object
Goal
Long
Avg.
Avg.
Len-5
π0
SFT
65.3
64.4
49.8
51.2
57.6
50.8
57.5
Flow-SDE
98.4
99.4
96.2
90.2
96.1
78.1
61.7
Flow-Noise
99.0
99.2
98.2
93.8
97.6
85.8
59.9
Uni-TMPO (ours)
99.2
99.6
98.8
93.6
97.8
88.6
63.9
π0.5
SFT
84.6
95.4
84.6
43.9
77.1
43.8
61.3
TABLE II: In-distribution VLA performance of π0 and π0.5 . Baseline configurations follow πRL [ 4 ] . LIBERO and MetaWorld-MT50 report average task success, while CALVIN-D reports five-subtask sequence completion (Len-5). All values are percentages. Bold values indicate the best result for each backbone.
Fig. 5: Task and scene OOD generalization after VLA post-training. MetaWorld ML45 uses 45 tasks for online RL and evaluates five unseen tasks. CALVIN uses Scenes ABC for supervised initialization and online RL, then evaluates five-subtask completion (Len-5) in Scene D over 1,000 sequences. The lower panels compare SFT and post-RL performance for π0 and π0.5 .
Fig. 6: WallDetour experiment. LEFT is shorter than RIGHT, although both reach the target under the same 0.9×Success+0.1×PathEfficiency reward. Reward Maximization concentrates successful rollouts on LEFT ( 0.95/0.05 ), whereas Uni-TMPO retains both routes ( 0.55/0.45 ), with normalized binary route entropies of 0.29 and 0.99 . When LEFT is blocked at evaluation, only Uni-TMPO reaches the target via RIGHT.
Fig. 7: Real-robot dual-target placement. The robot is instructed to place the red block on either yellow target. The left target receives slightly higher reward because it is closer. In the unblocked setting, both Reward Maximization rollouts select the left target, whereas Uni-TMPO succeeds at both targets. After the left target is blocked, Reward Maximization fails, while Uni-TMPO succeeds through the alternative right-target strategy.
Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from significant reward hacking, which degrades generative diversity and quality by inducing visual mode collapse and amplifying unreliable rewards. We identify the root cause as the mode-seeking nature of these methods, which maximize expected reward without effectively constraining probability distribution over acceptable trajectories, causing concentration on a few high-reward paths. In contrast, we propose Trajectory Matching Policy Optimization (TMPO), which replaces scalar reward maximization with trajectory-level reward distribution matching. Specifically, TMPO introduces a Softmax Trajectory Balance (Softmax-TB) objective to match the policy probabilities of K trajectories to a reward-induced Boltzmann distribution. We prove that this objective inherits the mode-covering property of forward KL divergence, preserving coverage over all acceptable trajectories while optimizing reward. To further reduce multi-trajectory training time on large-scale flow-matching models, TMPO incorporates Dynamic Stochastic Tree Sampling, where trajectories share denoising prefixes and branch at dynamically scheduled steps, reducing redundant computation while improving training effectiveness. Extensive results across diverse alignment tasks such as human preference, compositional generation and text rendering show that TMPO improves generative diversity over state-of-the-art methods by 9.1%, and achieves competitive performance in all downstream and efficiency metrics, attaining the optimal trade-off between reward and diversity.
Jiaming Li, Chenyu Zhu, Nanxi Yi +9
1MAIR Lab, Huazhong University of Science and Technology · 2Kuaishou Technology · 3Nanyang Technological University +2
Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the entire generative trajectory. However, text-to-image generation naturally has temporal and spatial structure: different denoising steps are responsible for different generation stages, and the content that truly determines text alignment often appears only in part of the image. This granularity mismatch makes it difficult for policy updates to focus on the generative components that actually affect the reward. To address this issue, we propose \textbf{SpatioTemporal Adaptive Reward (STAR) Allocation} for RL post-training of text-to-image diffusion and flow models. STAR uses text-image attention inside the generative model and starts from the core content that the user truly cares about in the prompt. It constructs spatial allocation maps that dynamically vary across denoising steps and rollouts, and allocates the same group-relative advantage to more relevant latent regions with almost no additional computational overhead. STAR then applies stronger policy updates to these regions through a spatially resolved policy objective. We use Stable Diffusion 3.5 Medium as the base model and evaluate on three tasks: GenEval, OCR text rendering, and PickScore. Experimental results show that STAR improves compositional semantic alignment, text rendering, and preference optimization without changing the external reward source, achieving 0.9759, 0.9757, and 23.60 on GenEval, OCR, and PickScore, respectively.
Reinforcement learning (RL) has become a standard technique for post-training diffusion-based image synthesis models, as it enables learning from reward signals to explicitly improve desirable aspects such as image quality and prompt alignment. In this paper, we propose an online RL variant that reduces the variance in the model updates by sampling paired trajectories and pulling the flow velocity in the direction of the more favorable image. Unlike existing methods that treat each sampling step as a separate policy action, we consider the entire sampling process as a single action. We experiment with both high-quality vision language models and off-the-shelf quality metrics for rewards, and evaluate the outputs using a broad set of metrics. Our method converges faster and yields higher output quality and prompt alignment than previous approaches.