HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training
Authors: Xinrui Chen, Mengyang Li, Ou Wu, Ji Zhang
Organizations: Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou, China · Tianjin Key Laboratory of Wireless Mobile Communications and Power Transmission, Tianjin Normal University, Tianjin, China · University of Southern Queensland
Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-management methods set state fidelity from execution cost, tensor properties, or generic compression sensitivity, without explicitly incorporating GRPO's analytic update structure into state-fidelity allocation. We formalize this dependence as policy-update exposure, linking the current GRPO loss coefficients to state-level approximation sensitivity. These coefficients are available before backward without an additional backward pass. We introduce HiLoRe, which allocates graph-attributed recovery units among high-precision storage, low-precision compression, and deterministic recomputation using measured recovery utility and update-conditioned approximation risk. It combines high-precision storage and deterministic recomputation with low-precision recovery under a calibrated risk budget. Across five model-task settings with 2K responses and memory < 1.10 times GC's per-GPU actor-update peak, HiLoRe's actor-update throughput gains reach 13.5% over GC and 7.9% over the fastest evaluated baseline, with paired mean downstream-score differences below 0.6 percentage points.
Figures & tables
Figure 1: Motivation of update-conditioned recovery. Fixed recovery underutilizes memory headroom, while GRPO exposes update-dependent state fidelity before backward.
Figure 2: HiLoRe maps GRPO-native signals and utility to fidelity-constrained H/L/R scheduling.
Algorithm 1 HiLoRe Recovery Allocation
Model
Method
Efficiency
Gradient Fidelity
Training Quality
Peak MiB ↓
Tok./s ↑
Gain (%) ↑
Grad. err. ↓
MATH500 ↑
GSM8K ↑
Qwen2.5-3B
GC / All-R \venuebox arXiv’16
28,874
2,455.72
–
0.0000
60.20±1.00
77.33±0.15
Rockmate \venuebox ICML’23
29,992
2,066.73
−15.84±0.79
0.0031
60.20±0.80
76.90±0.04
ALAM \venuebox ICLR’24
28,595
2,381.56
−3.02±0.57
0.0131
59.13±0.50
76.57±0.23
Adacc \venuebox arXiv’25
31,400
2,579.00
+5.02±0.49
0.0107
60.00±0.40
77.08±0.12
INSTANT \venuebox ICLR’26
29,236
2,515.39
+2.43±0.47
0.0105
59.80±0.40
77.00±0.12
Table 1: Actor-update efficiency, gradient fidelity, and downstream quality on DeepMath10K with 2K responses. Red and orange indicate the best and second-best values within each model.
Figure 3: Actor-update time breakdown.
Figure 6
Figure 5: Memory–throughput frontier, update fidelity, and recovery allocation.
Table 8
Figure 6: Update sparsity, risk prediction, and component ablations in post-update diagnostic replays.
Update regime
Norm. exposure
H (%)
L (%)
Low ∣At∣
0.48
4.8
18.5
High ∣At∣
1.71
14.7
9.4
Clipped
0.27
2.9
21.3
Active unclipped
1.43
12.6
11.2
Table 5: Exposure and recovery allocation.
Figure 7: Sensitivity to the internal fidelity budget δ , calibration size, and risk-refresh interval.
Table 13: Recovery configurations under memory ceilings.
Figure 8: Response-length and model transfer. Throughput gains are relative to paired GC, with error bars showing sample SDs. Gradient errors and additional peak memory are expressed as percentages.
Setting
Δ Metric 1
Δ Metric 2
TACO-Verified / Qwen2.5-3B
−0.06±0.36
+0.09±0.40
Logic-RL K&K / Qwen2.5-3B
+0.07±0.31
+0.10±0.17
Phi-3.5-mini / DeepMath10K
+0.53±0.50
+0.13±0.19
Llama-3.1-8B / DeepMath10K
+0.53±0.46
+0.23±0.23
Appendix
Table 14: Paired downstream quality changes relative to GC, in percentage points.
Figure 9: Full-parameter memory–throughput trade-offs. GC (tuned) adjusts microbatch size per budget with decoder-layer checkpointing.
ϵg
HiLoRe
Adacc
AGoQ
PRAC
INSTANT
1.00%
8.76 / 0.78
4.30 / 0.95
2.55 / 0.93
1.85 / 0.94
1.70 / 0.96
1.25%
10.16 / 1.04
5.02 / 1.07
3.18 / 1.10
2.76 / 1.12
2.43 / 1.05
1.50%
10.16 / 1.04
5.02 / 1.07
3.18 / 1.10
2.76 / 1.12
2.43 / 1.05
1.75%
10.94 / 1.58
7.10 / 1.68
6.20 / 1.73
6.70 / 1.65
7.40 / 1.67
2.00%
11.28 / 1.81
8.10 / 1.93
7.20 / 1.96
8.70 / 1.94
9.70 / 1.91
Appendix
Table 15: Shared gradient-error tolerance sensitivity. Each cell reports throughput gain over GC (%) / mean gradient error (%).
Regime
Norm. ∣ωt∣
Norm. exposure
H (%)
L (%)
L drift
Low ∣At∣
0.41
0.48
4.8
18.5
0.0062
High ∣At∣
1.84
1.71
14.7
9.4
0.0127
Clipped
0.19
0.27
2.9
21.3
0.0051
Active unclipped
1.37
1.43
12.6
11.2
0.0115
Appendix
Table 16: Update-regime statistics averaged over held-out post-update diagnostic replays.
Sweep
Setting
Gain (%) ↑
Mean err. ↓
Max err. ↓
Scheduling overhead (%)
H/L/R (%)
δ/δ⋆
0
+6.14±0.48
0.0030
0.0034
0.19
13.9/0/86.1
0.25
+7.35±0.47
0.0057
0.0064
0.21
11.1/5.6/83.3
0.5
+8.76±0.46
0.0078
0.0088
0.21
8.3/11.1/80.6
1
+10.16±0.44
0.0104
0.0114
0.21
8.3/16.7/75.0
2
+10.94±0.46
0.0158
0.0176
0.21
5.6/22.2/72.2
4
+11.28±0.49
0.0181
0.0206
0.22
2.8/30.6/66.6
Appendix
Table 17: Sensitivity to risk budget, calibration size, and refresh interval. Scheduling overhead includes exposure/risk refresh and allocation.
Component
GC
HiLoRe-HR
HiLoRe
Forward
146.72
146.85
146.91
Recomputation
207.38
164.10
118.64
Candidate encoding
–
–
6.37
L reconstruction
–
–
7.88
Exposure refresh
–
0.72
0.74
Allocation
–
0.51
0.58
Appendix
Table 18: Profiled actor-update components in seconds; totals sum the listed components.
Department of Electrical and Computer Engineering, Seoul National University · Interdisciplinary Program in AI, Seoul National University · AIIS, ASRI, INMC, and ISRC, Seoul National University