Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Representative DPO-, GRPO-, and NFT-style objectives, when applied to negative trajectories, mainly steer predicted velocities away from the corresponding directions without explicitly specifying a target velocity field toward preferred samples. In 3D generation, constrained by pretrained model capabilities, rollout diversity, and reward-distribution complexity, directly applying these RL methods yields limited gains in geometric quality. We introduce a forward-process RL method \textbf{Dynamic Homing Optimization (DHO)}, which reformulates negative-trajectory optimization as positive-sample attraction-guided dynamic homing. Specifically, Minimum-Cost Attractive Matching (MAM) assigns each negative sample a distinct positive target, and Time-Aware Dynamic Correction (TDC) then redirects its trajectory toward the target using a remaining-time-aware corrective velocity. Building on asynchronous online DHO, we develop \textbf{Flow3D-Pro}, an image-to-3D geometry generation framework. Experiments show that DHO outperforms representative DPO-, GRPO-, and NFT-style objectives in 3D generation, while Flow3D-Pro produces higher-quality 3D geometry than existing mesh generation methods.
Figures & tables
Figure 1: We generate diverse 3D assets with plausible global shapes and detailed geometry.
Figure 2: Negative-sample branch objectives across flow matching RL methods.
Figure 3: Method Overview. After pre-training and supervised fine-tuning (SFT), DHO post-trains the rectified flow model in latent feature space using reward-ranked rollout samples. In DHO, MAM assigns each negative sample an attractive positive-target destination, and TDC then constructs a time-aware homing velocity toward the assigned target. We visualize only the negative-sample branch, as the positive-sample branch uses the same MAM rule and homing-velocity construction.
Figure 4: The generated owl gradually becomes canonically aligned during DHO training.
Method
Stage I
Stage II
ULIP ↑
Uni3D ↑
User Study ↑
ULIP ↑
Uni3D ↑
User Study ↑
Base
0.143
0.347
0.60
0.143
0.350
0.71
SFT
0.144 (+0.70%)
0.351 (+1.15%)
1.93
0.143 (+0.00%)
0.352 (+0.57%)
1.91
GRPO-style
0.147 (+2.80%)
0.357 (+2.88%)
3.02
0.147 (+2.80%)
0.358 (+2.29%)
3.00
DPO-style
0.146 (+2.10%)
0.357 (+2.88%)
2.36
0.146 (+2.10%)
0.358 (+2.29%)
2.32
NFT-style
0.146 (+2.10%)
0.355 (+2.31%)
3.07
0.147 (+2.80%)
0.357 (+2.00%)
2.80
Table 1: Comparison of RL objectives for Stage I, evaluated before and after the Stage II refiner.
Figure 5: Qualitative comparison of RL objectives. To facilitate visualization, all generated meshes are rotated to a common orientation. The same convention is used in subsequent visualizations.
Figure 6: Visual comparison with representative 3D generation methods.
Method
ULIP ↑
Uni3D ↑
User Study ↑
Michelangelo Zhao et al. (2023)
0.135
0.290
0.18
Craftsman 1.5 Li et al. (2025b)
0.146
0.325
1.32
Hi3DGen Ye et al. (2025)
0.144
0.335
3.77
Direct3D-s2 Wu et al. (2025b)
0.141
0.345
3.27
Hunyuan3D 2.1 Yang et al. (2025)
0.143
0.347
4.14
Trellis Xiang et al. (2025)
0.145
0.351
4.27
Table 2: Quantitative evaluation on 3D mesh generation.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Illustration of the VLM-based reward evaluation process. The reward model receives a reference image, six canonical-view renderings of the generated mesh, and an evaluation prompt template, then produces criterion-wise assessments and a final scalar reward.
Figure 8: Effect of the target-assignment strategy in DHO.
Figure 9: Effect of positive-sample selection in MAM.
Time Weighting
ULIP ↑
Uni3D ↑
User Study ↑
w(t)=1
0.146
0.354
0.45
w(t)=(1−t)2
0.149 (+2.05%)
0.359 (+1.41%)
0.55
Appendix
Table 3: Effect of time-aware weighting in TDC.
Velocity Target
ULIP ↑
Uni3D ↑
User Study ↑
Static
0.145
0.353
0.31
Dynamic
0.149 (+2.76%)
0.359 (+1.70%)
0.69
Appendix
Table 4: Static versus dynamic velocity targets for the negative branch.
Advantage Weighting
ULIP ↑
Uni3D ↑
User Study ↑
Without
0.145
0.354
0.42
With
0.149 (+2.76%)
0.359 (+1.41%)
0.58
Appendix
Table 5: Ablation of positive-sample advantage weighting in DHO.
Figure 10: Ablation study on the CFG scale.
η
β
ULIP ↑
Uni3D ↑
User Study ↑
0.0
0.1
0.2
0.5
0.0
1.0
5.0
✓
✓
0.145
0.354
1.13
✓
✓
0.149 (+2.76%)
0.359 (+1.41%)
1.92
✓
✓
0.147 (+1.38%)
0.357 (+0.85%)
1.68
✓
✓
0.145 (+0.00%)
0.355 (+0.28%)
1.27
✓
✓
0.140
0.342
0.30
Appendix
Table 6: DHO performance with the homing coefficient η and the regularization coefficient β . The upper block fixes β and varies η (where η=0.0 disables the dynamic homing branch for negative-samples), while the lower block fixes η and varies β (where β=0.0 removes the reference-policy regularization). Each block is an independent User Study comparison group.
Figure 11: Effect of the rollout group size G on DHO performance.
Rollout Sampler
ULIP ↑
Uni3D ↑
User Study ↑
ODE
0.145
0.353
0.33
SDE
0.149 (+2.76%)
0.359 (+1.70%)
0.67
Appendix
Table 7: Comparison of ODE and SDE rollout samplers.
Figure 12: DHO training dynamics with ODE and SDE rollouts. From left to right: negative-branch gain, positive-branch gain, and reference-policy deviation.
Reconstruction Fidelity
Geometric Plausibility
Orientation Alignment
User Study ↑
–
–
–
0.63
✓
1.46
✓
✓
1.85
✓
✓
✓
2.06
Appendix
Table 8: Impact of the reward dimensions used for DHO post-training, where “–” denotes the SFT model.
Method
ULIP ↑
Uni3D ↑
Michelangelo Zhao et al. (2023)
0.115
0.261
Craftsman 1.5 Li et al. (2025b)
0.129
0.298
Trellis Xiang et al. (2025)
0.126
0.311
Hunyuan3D 2.0 Zhao et al. (2025)
0.130
0.315
Hi3DGen Ye et al. (2025)
0.112
0.299
Direct3D-s2 Wu et al. (2025b)
0.122
0.314
Appendix
Table 9: Quantitative evaluation on the LATTICE-Bench.
Method
ULIP ↑
Uni3D ↑
User Study ↑
SFT
0.144
0.351
0.97
RFT
0.145 (+0.69%)
0.355 (+1.14%)
1.39
PFM Kim et al. (2024)
0.147 (+2.08%)
0.356 (+1.42%)
1.70
DHO
0.149 (+3.47%)
0.359 (+2.28%)
1.94
Appendix
Table 10: Comparison with additional post-training methods.
Figure 13: Visual comparison against commercial models. To facilitate visualization, all generated meshes are rotated to a common orientation.
Figure 14: More qualitative results of Flow3D-Pro across diverse object categories and geometric structures.