cs.LGOct 8, 2026

How to post-train on a surrogate: Envelope sampling mitigates reward hacking

Authors: Sanjit Dandapanthula, Shuvom Sadhuka, Samir Khan, Michael Oberst, Aaditya Ramdas, Alexandra Chouldechova

Abstract

Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale. This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects. In this work, we study a setting in which a small number nn of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it. Prior approaches to judge recalibration are costly or heuristic, and it is known that on-policy sampling fails when the surrogate is miscalibrated on a rare set of outputs. In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an L2L^2 ball around the judge. We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward, and experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.

Explore similar work

CardsList
  1. Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR

    Date pendingMuhammad Khalifa, Zohaib Khan, Omer Tafveez +2RL BenchmarksReward Hacking

  2. Small Reward Models via Backward Inference

    Feb 14, 2026Yike Wang, Faeze Brahman, Shangbin Feng +3Reward ModelingRL for Language Models

  3. Sharpening Tax in Post-Training

    Oct 1, 2026Changdae Oh, Qi Zeng, Qi Qi +7Agentic RLTest-Time Scaling