HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training
Authors: Xinrui Chen, Mengyang Li, Ou Wu, Ji Zhang
Organizations: Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou, China · Tianjin Key Laboratory of Wireless Mobile Communications and Power Transmission, Tianjin Normal University, Tianjin, China · University of Southern Queensland
Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-management methods set state fidelity from execution cost, tensor properties, or generic compression sensitivity, without explicitly incorporating GRPO's analytic update structure into state-fidelity allocation. We formalize this dependence as policy-update exposure, linking the current GRPO loss coefficients to state-level approximation sensitivity. These coefficients are available before backward without an additional backward pass. We introduce HiLoRe, which allocates graph-attributed recovery units among high-precision storage, low-precision compression, and deterministic recomputation using measured recovery utility and update-conditioned approximation risk. It combines high-precision storage and deterministic recomputation with low-precision recovery under a calibrated risk budget. Across five model-task settings with 2K responses and memory < 1.10 times GC's per-GPU actor-update peak, HiLoRe's actor-update throughput gains reach 13.5% over GC and 7.9% over the fastest evaluated baseline, with paired mean downstream-score differences below 0.6 percentage points.
Figures & tables
Figure 1: Motivation of update-conditioned recovery. Fixed recovery underutilizes memory headroom, while GRPO exposes update-dependent state fidelity before backward.
Figure 2: HiLoRe maps GRPO-native signals and utility to fidelity-constrained H/L/R scheduling.
Algorithm 1 HiLoRe Recovery Allocation
Model
Method
Efficiency
Gradient Fidelity
Training Quality
Peak MiB ↓
Tok./s ↑
Gain (%) ↑
Grad. err. ↓
MATH500 ↑
GSM8K ↑
Qwen2.5-3B
GC / All-R \venuebox arXiv’16
28,874
2,455.72
–
0.0000
60.20±1.00
77.33±0.15
Rockmate \venuebox ICML’23
29,992
2,066.73
−15.84±0.79
0.0031
60.20±0.80
76.90±0.04
ALAM \venuebox ICLR’24
28,595
2,381.56
−3.02±0.57
0.0131
59.13±0.50
76.57±0.23
Adacc \venuebox arXiv’25
31,400
2,579.00
+5.02±0.49
0.0107
60.00±0.40
77.08±0.12
INSTANT \venuebox ICLR’26
29,236
2,515.39
+2.43±0.47
0.0105
59.80±0.40
77.00±0.12
Table 1: Actor-update efficiency, gradient fidelity, and downstream quality on DeepMath10K with 2K responses. Red and orange indicate the best and second-best values within each model.
Figure 3: Actor-update time breakdown.
Figure 6
Figure 5: Memory–throughput frontier, update fidelity, and recovery allocation.
Table 8
Figure 6: Update sparsity, risk prediction, and component ablations in post-update diagnostic replays.
Update regime
Norm. exposure
H (%)
L (%)
Low ∣At∣
0.48
4.8
18.5
High ∣At∣
1.71
14.7
9.4
Clipped
0.27
2.9
21.3
Active unclipped
1.43
12.6
11.2
Table 5: Exposure and recovery allocation.
Figure 7: Sensitivity to the internal fidelity budget δ , calibration size, and risk-refresh interval.
Table 13: Recovery configurations under memory ceilings.
Figure 8: Response-length and model transfer. Throughput gains are relative to paired GC, with error bars showing sample SDs. Gradient errors and additional peak memory are expressed as percentages.
Setting
Δ Metric 1
Δ Metric 2
TACO-Verified / Qwen2.5-3B
−0.06±0.36
+0.09±0.40
Logic-RL K&K / Qwen2.5-3B
+0.07±0.31
+0.10±0.17
Phi-3.5-mini / DeepMath10K
+0.53±0.50
+0.13±0.19
Llama-3.1-8B / DeepMath10K
+0.53±0.46
+0.23±0.23
Appendix
Table 14: Paired downstream quality changes relative to GC, in percentage points.
Figure 9: Full-parameter memory–throughput trade-offs. GC (tuned) adjusts microbatch size per budget with decoder-layer checkpointing.
ϵg
HiLoRe
Adacc
AGoQ
PRAC
INSTANT
1.00%
8.76 / 0.78
4.30 / 0.95
2.55 / 0.93
1.85 / 0.94
1.70 / 0.96
1.25%
10.16 / 1.04
5.02 / 1.07
3.18 / 1.10
2.76 / 1.12
2.43 / 1.05
1.50%
10.16 / 1.04
5.02 / 1.07
3.18 / 1.10
2.76 / 1.12
2.43 / 1.05
1.75%
10.94 / 1.58
7.10 / 1.68
6.20 / 1.73
6.70 / 1.65
7.40 / 1.67
2.00%
11.28 / 1.81
8.10 / 1.93
7.20 / 1.96
8.70 / 1.94
9.70 / 1.91
Appendix
Table 15: Shared gradient-error tolerance sensitivity. Each cell reports throughput gain over GC (%) / mean gradient error (%).
Regime
Norm. ∣ωt∣
Norm. exposure
H (%)
L (%)
L drift
Low ∣At∣
0.41
0.48
4.8
18.5
0.0062
High ∣At∣
1.84
1.71
14.7
9.4
0.0127
Clipped
0.19
0.27
2.9
21.3
0.0051
Active unclipped
1.37
1.43
12.6
11.2
0.0115
Appendix
Table 16: Update-regime statistics averaged over held-out post-update diagnostic replays.
Sweep
Setting
Gain (%) ↑
Mean err. ↓
Max err. ↓
Scheduling overhead (%)
H/L/R (%)
δ/δ⋆
0
+6.14±0.48
0.0030
0.0034
0.19
13.9/0/86.1
0.25
+7.35±0.47
0.0057
0.0064
0.21
11.1/5.6/83.3
0.5
+8.76±0.46
0.0078
0.0088
0.21
8.3/11.1/80.6
1
+10.16±0.44
0.0104
0.0114
0.21
8.3/16.7/75.0
2
+10.94±0.46
0.0158
0.0176
0.21
5.6/22.2/72.2
4
+11.28±0.49
0.0181
0.0206
0.22
2.8/30.6/66.6
Appendix
Table 17: Sensitivity to risk budget, calibration size, and refresh interval. Scheduling overhead includes exposure/risk refresh and allocation.
Component
GC
HiLoRe-HR
HiLoRe
Forward
146.72
146.85
146.91
Recomputation
207.38
164.10
118.64
Candidate encoding
–
–
6.37
L reconstruction
–
–
7.88
Exposure refresh
–
0.72
0.74
Allocation
–
0.51
0.58
Appendix
Table 18: Profiled actor-update components in seconds; totals sum the listed components.
Reinforcement learning from verifiable rewards with GRPO is a standard approach for post-training reasoning LLMs. It remains sample inefficient. Each rollout is used for a single gradient update and then discarded. Naive replay is not well suited in this setting because LLM policies drift quickly per gradient step. Stored rollouts therefore become stale and can destabilize training. We propose a rollout-level replay buffer for GRPO that stores and samples individual rollouts rather than whole groups. The buffer bounds staleness through age eviction. Any rollout older than tau_max training steps is removed. The buffer also preserves on-policy data via fresh-anchored composition. Each batch keeps its fresh on-policy rollouts and then concatenates replay rollouts drawn separately from the buffer. We prioritize replay by per-rollout advantage magnitude and recycle individual rollouts whose advantages are large. Across three Qwen3-Base scales on five math benchmarks, our method outperforms GRPO and naive replay baselines. Gains are positive at every scale and reach +1.66 pp on the five-benchmark average at 4B. Under an AES metric that jointly measures accuracy and token efficiency, our method is the only condition with a positive margin over GRPO at every scale.
Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang +2
Department of Electrical and Computer Engineering, Seoul National University · Interdisciplinary Program in AI, Seoul National University · AIIS, ASRI, INMC, and ISRC, Seoul National University
Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. Across reasoning tasks, LoGRA reduces average training memory by up to 45.7% without sacrificing performance. It also enables stable training of a 27B-parameter model for over 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory, making previously memory-infeasible RL training practical. Code is available in the Molt library.
Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including DanceGRPO and FlowGRPO, recompute selected timesteps with gradient tracking after rollout. Under on-policy training with the same backend for rollout and update, this recomputation is mathematically redundant. Intuitively, the rollout and policy update steps can reuse the same feed-forward backbone to avoid redundant computation, but doing so can incur a large memory overhead during rollout. To address the issue, we present LeanGRPO by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL: (1) LeanGRPO-Retain enables gradient tracking during rollout and directly reuses the resulting computation graphs and saved activations for backward during update, requiring no recomputation; and (2) LeanGRPO-Reweight also enables gradients during rollout, but immediately backpropagates each selected step using a provisional advantage and delays gradient synchronization, then corrects the provisional gradients with the true advantage after the trajectory is completed. These schedules target different model scales and input sizes. Across FlowGRPO/DanceGRPO with FLUX.1-dev and Wan, LeanGRPO achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.
Sijie Wang, Zhiqiang Tan, Xinrui Yang +1
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen