Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization
Organizations: School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China
Abstract
Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.
Figures & tables
| Method | Time (s) | GenEval | OCR | PickScore | HPS | ImgRwd | LGMD | Cos.Div. |
|---|---|---|---|---|---|---|---|---|
| Compositional Image Generation (GenEval) | ||||||||
| FLUX.1-dev | – | 0.647 | – | 22.301 | 0.301 | 1.099 | 0.211 | |
| DAG-DB | 187.5 | 0.889 | – | 21.998 | 0.291 | 1.071 | 0.097 | 0.237 |
| DGFS-SubTB | 178.6 | 0.917 | – | 22.210 | 0.298 | 1.107 | 0.113 | 0.241 |
| Flow-GRPO | 160.8 | 0.946 | – | 22.113 | 0.289 | 1.074 | 0.198 | |
| TreeGRPO | 126.2 | 0.936 | – | 21.524 | 0.281 | 1.083 | 0.184 | |
| Backbone | Method | LIBERO | MetaWorld-MT50 | CALVIN-D | ||||
|---|---|---|---|---|---|---|---|---|
| Spatial | Object | Goal | Long | Avg. | Avg. | Len-5 | ||
| SFT | 65.3 | 64.4 | 49.8 | 51.2 | 57.6 | 50.8 | 57.5 | |
| Flow-SDE | 98.4 | 99.4 | 96.2 | 90.2 | 96.1 | 78.1 | 61.7 | |
| Flow-Noise | 99.0 | 99.2 | 98.2 | 93.8 | 97.6 | 85.8 | 59.9 | |
| Uni-TMPO (ours) | 99.2 | 99.6 | 98.8 | 93.6 | 97.8 | 88.6 | 63.9 | |
| SFT | 84.6 | 95.4 | 84.6 | 43.9 | 77.1 | 43.8 | 61.3 | |