OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search
Organizations: Dobot Robotics · Independent Researcher · Zhejiang University · Fudan University · Osaka University
Abstract
The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while controlling gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Trajectory Policy Optimization (TTPO) using on-policy tree trajectories, which sample new suffixes from the current policy at visited states. This needs no action-distribution correction, although branching changes state visitation. Our Branch Aggregation Lemma shows that branch-weighted tree statistics recover chain expectations when branch choices and weights are fixed before outgoing transitions are sampled. OPTS selects expansion states using estimated performance differences. Under deterministic dynamics, exact values, and max-backup advantages, the induced search policy's expected return improves monotonically with the budget. We bound the gradient bias from adaptive expansion and show that max backup assigns prefix credit to actions leading to better discovered suffixes. Against a finite chain reference, TTPG's measured bias stays near its no-branching level, while NaivePG's bias grows from 0.1251 to 0.4884. At matched budgets, reward- and value-guided OPTS improve correct-answer coverage and majority-vote accuracy over independent sampling. At matched branch counts, OPTS + TTPG gains coverage with a modest bias increase relative to Fixed-branch + TTPG. Under matched interaction or rollout budgets, OPTS-TTPO improves MuJoCo tail returns over PPO by up to 28.6%, achieves a 34-22-1 win-loss-tie record against PPO on Atari-57 under the last-100-log mean-return metric, and improves micro-averaged avg@32 and pass@32 over PPO across all four Qwen3 models.
Figures & tables
| Variant | Hopper | Walker2d | HalfCheetah | Ant | Humanoid | Mean |
|---|---|---|---|---|---|---|
| Random + mean | +10.5 | +5.6 | -3.2 | +15.0 | -1.3 | +5.3 |
| Guided + mean | +8.2 | +16.7 | +9.5 | -0.6 | +6.3 | +8.0 |
| Guided + max, unweighted | +7.5 | +6.8 | +26.2 | +8.7 | +8.2 | +11.5 |
| Guided + max (full) | +11.2 | +24.9 | +28.6 | +13.9 | +15.0 | +18.7 |
| Human-normalized IQM | Task wins | ||||
|---|---|---|---|---|---|
| Metric | PPO | OPTS-TTPO | PPO | OPTS-TTPO | Tie |
| Full-training mean return | 0.247 | 0.255 | 26 | 31 | 0 |
| Last-100-log mean return | 0.357 | 0.374 | 22 | 34 | 1 |
| Methods | MATH500 | MinervaMath | AMC23 | AIME24 | AIME25 | AIME26 | Macro Average | Micro Average | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| avg@32 | pass@32 | avg@32 | pass@32 | avg@32 | pass@32 | avg@32 | pass@32 | avg@32 | pass@32 | avg@32 | pass@32 | avg@32 | pass@32 | avg@32 | pass@32 | ||
| Qwen3-1.7B-Base | |||||||||||||||||
| PPO | 0.6973 | 0.9080 | 0.2986 | 0.5515 | 0.4289 | 0.8500 | 0.0823 | 0.3333 | 0.0469 | 0.3333 | 0.0417 | 0.2667 | 0.2659 | 0.5405 | 0.5013 | 0.7384 | |
| DAPO | 0.6935 | 0.9100 | 0.2878 | 0.5257 | 0.4133 | 0.8750 | 0.0854 | 0.3667 | 0.0490 | 0.3000 | 0.0385 | 0.2667 | 0.2612 | 0.5407 | 0.4953 | 0.7328 | |
| REINFORCE++ | 0.6877 | 0.9040 | 0.2911 | 0.5184 | 0.4016 | 0.8250 | 0.0823 | 0.3000 | 0.0521 | 0.3667 | 0.0354 | 0.2667 | 0.2584 | 0.5301 | 0.4924 | 0.7251 | |
| OPTS-TTPO | 0.7114 | 0.9220 | 0.2920 | 0.5404 | 0.4328 | 0.8000 | 0.1063 | 0.4333 | 0.0479 | 0.3000 | 0.0521 | 0.2667 | 0.2738 | 0.5437 | 0.5085 | 0.7428 | |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | MuJoCo | Atari-57 |
|---|---|---|
| Optimizer | Adam, | Adam, |
| Learning rate | , linearly annealed to 0 | , linearly annealed to 0 |
| Rollout batch | transitions | transitions |
| Minibatches / update epochs | 32 (64 transitions each) / 10 | 4 (256 transitions each) / 4 |
| Discount / GAE parameter | , | , |
| Policy clip / clipped value loss | 0.2 / yes | 0.1 / yes |
| Method | Method-specific settings |
|---|---|
| PPO | 4,096 prompts and rollouts per update; GAE with and ; learned critic with learning rate and value clip 0.5; symmetric policy clip 0.2; token-level loss aggregation. |
| DAPO | 512 prompts with 8 responses each; group-relative advantages and no critic; asymmetric policy clip with dual-clip constant 10; token-level loss aggregation; overlong-response buffer of 1,024 tokens with penalty factor 1.0. |
| REINFORCE++ | 512 prompts with 8 responses each; per-prompt mean reward baseline followed by global advantage whitening, with no critic; symmetric policy clip 0.2; token-level loss aggregation. |
| OPTS-TTPO | 4,096 rollouts collected over the search rounds specified in Appendix D ; TreeGAE with a learned critic at learning rate ; symmetric policy clip 0.2; branch-weighted token-level loss aggregation. |
| Metric | Comparison | Wins | Losses | Ties |
|---|---|---|---|---|
| Full training | Max vs. Mean | 19 | 22 | 16 |
| Full training | Max vs. PPO | 27 | 30 | 0 |
| Full training | Mean vs. PPO | 31 | 26 | 0 |
| Tail | Max vs. Mean | 18 | 22 | 17 |
| Tail | Max vs. PPO | 34 | 22 | 1 |
| Tail | Mean vs. PPO | 34 | 22 | 1 |
| Full | Tail | |||||
|---|---|---|---|---|---|---|
| Game | PPO | Mean | Max | PPO | Mean | Max |
| ALE_Surround-v5 | -2.83 | -2.94 | -2.94 | -1.14 | -0.91 | -0.91 |
| AlienNoFrameskip-v4 | 989.13 | 947.58 | 945.15 | 1186.07 | 1236.76 | 1372.65 |
| AmidarNoFrameskip-v4 | 173.34 | 192.09 | 191.86 | 246.73 | 301.20 | 307.77 |
| AssaultNoFrameskip-v4 | 922.34 | 967.38 | 968.90 | 1073.07 | 1278.37 | 1208.08 |
| AsterixNoFrameskip-v4 | 1553.26 | 1671.41 | 1690.02 | 2051.14 | 2223.03 | 2194.67 |
| Full | Tail | |||||
|---|---|---|---|---|---|---|
| Game | PPO | Mean | Max | PPO | Mean | Max |
| KangarooNoFrameskip-v4 | 1229.65 | 1612.89 | 1587.39 | 1802.94 | 2458.78 | 2444.39 |
| KrullNoFrameskip-v4 | 3762.75 | 3802.40 | 3776.63 | 4173.87 | 4482.65 | 4402.41 |
| KungFuMasterNoFrameskip-v4 | 6102.91 | 5726.26 | 5726.26 | 5609.22 | 6604.17 | 6604.17 |
| MontezumaRevengeNoFrameskip-v4 | 0.21 | 0.06 | 0.48 | 0.00 | 0.17 | 0.00 |
| MsPacmanNoFrameskip-v4 | 1185.86 | 1131.68 | 1184.35 | 1471.88 | 1464.83 | 1551.24 |
| Stage of one training step | Seconds | Share |
|---|---|---|
| Rollout generation (vLLM, all rounds) | 348.1 | 43.7% |
| Critic update | 214.4 | 26.9% |
| Actor update | 131.4 | 16.5% |
| Critic value forwards (all rounds) | 56.5 | 7.1% |
| Old log-probabilities | 39.2 | 4.9% |
| Search module: path refresh + selection | 3.2 | 0.4% |