Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage. We ask how to make process supervision efficient: accelerating convergence and improving final quality without the cost of value networks. We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics. BoT-GRPO is critic-free, and is a drop-in replacement wherever GRPO is used when token-level reward is available. On React front-end code generation, BoT-GRPO reaches 80% compile rate up to 1.9× faster than GRPO and converges faster than modern GRPO variants (GSPO, DAPO, PURE) while reaching higher final compile and VLM-judged win rates. On a second task, AIME mathematical reasoning, BoT-GRPO delivers absolute Pass@k gains up to 8.1% over GRPO in half the steps. For both tasks we compare the algorithm's performance on reasoning vs. non-reasoning base-model families (Qwen2.5-3B, SmolLM3-3B, Phi-4-mini-reasoning). Our experiments also yield a practical recipe for the reward model itself: reward stability matters more than richness: clean, bounded, stable fine-grained signals consistently accelerate learning where noisier alternatives stall.
Figures & tables
Figure 1: React compile success rate on the held-out set for BoT-GRPO and four RL baselines. BoT-GRPO (blue, dashed) is the first to cross every compile-rate threshold; PURE (red) matches early but degrades later, consistent with instability from unbounded process-reward aggregation.
Figure 2: Pairwise VLM win rate (4-round ABBA) of BoT-GRPO against each baseline (ties =0.5 ; BoT-GRPO above 0.5 means it is preferred; grey bars show the tie rate). BoT-GRPO wins throughout training against all four – most decisively over vanilla GRPO ( 62 – 90% ) and PURE ( 67 – 78% ), and by a durable margin over GSPO ( 56 – 78% ) and DAPO ( 58 – 63% ). The GSPO/DAPO gaps reflect visual quality improvements beyond mere compilability, since all three saturate on compile rate.
Figure 3: React compile-rate convergence for BoT-GRPO vs. GRPO across three base-model families (Qwen2.5-3B, SmolLM3-3B, Phi-4-mini-reasoning).
Figure 4: Effect of the reward stack on React (Qwen2.5-3B), validation compile rate over training. GRPO (global) and the two BoT-GRPO variants include the LLM code judge (the local variant adds the VLM screenshot signal); the red dashed run drops the judge and uses compiler + ESLint only .
Figure 5: BoT-GRPO vs. standard GRPO on AIME (Qwen2.5-3B).
Figure 6: BoT-GRPO vs. GRPO on AIME across three base models (Qwen2.5-3B, SmolLM3-3B, Phi-4-mini-reasoning).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Effect of local reward strategy on AIME. Error propagation enables effective learning; context-aware and independent evaluation produce unstable signals that prevent convergence.
Figure 8: Sensitivity to local reward weight ws on AIME. Performance is stable across ws∈{0.1,0.5,1.0} when the global reward dominates.
Figure 9: Effect of step evaluation window size on AIME. Windows of 1 and 3 perform comparably; window 5 shows slightly slower early convergence.
Figure 10: VLM reward stability: four configurations differing in VLM screenshot judge weight and debiasing. Shaded region highlights instability from strong local VLM reward. Bar chart: peak-to-final compile rate drop.