cs.LGSep 27, 2026

HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training

Authors: Xinrui Chen, Mengyang Li, Ou Wu, Ji Zhang

Organizations: Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou, China · Tianjin Key Laboratory of Wireless Mobile Communications and Power Transmission, Tianjin Normal University, Tianjin, China · University of Southern Queensland

Abstract

Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-management methods set state fidelity from execution cost, tensor properties, or generic compression sensitivity, without explicitly incorporating GRPO's analytic update structure into state-fidelity allocation. We formalize this dependence as policy-update exposure, linking the current GRPO loss coefficients to state-level approximation sensitivity. These coefficients are available before backward without an additional backward pass. We introduce HiLoRe, which allocates graph-attributed recovery units among high-precision storage, low-precision compression, and deterministic recomputation using measured recovery utility and update-conditioned approximation risk. It combines high-precision storage and deterministic recomputation with low-precision recovery under a calibrated risk budget. Across five model-task settings with 2K responses and memory < 1.10 times GC's per-GPU actor-update peak, HiLoRe's actor-update throughput gains reach 13.5% over GC and 7.9% over the fastest evaluated baseline, with paired mean downstream-score differences below 0.6 percentage points.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rollout-Level Advantage-Prioritized Experience Replay for GRPO

    Jun 3, 2026Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang +2Parallel RolloutsReplay

  2. LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

    Sep 3, 2026Sijie Wang, Zhiqiang Tan, Xinrui Yang +1Diffusion-Based Reinforcement Learning MethodsReward Gradients