cs.LGOct 5, 2026

Reward-Driven Learning under Prompt-Level Differential Privacy

Authors: Jiachen Zhao, Antonia Januszewicz, Taeho Jung

Organizations: University of Notre Dame

Abstract

Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be (ε,δ)-differentially private with respect to the presence of any one training problem. Taking the group of responses to one prompt as the privacy record, our method aggregates their gradients, clips the prompt's contribution once, adds Gaussian noise, and composes the privacy loss across updates, so the budget depends on neither the number of responses per prompt nor the clipping norm; to our knowledge this is the first differential privacy guarantee for RLVR training. We train Qwen2.5-1.5B-Instruct with LoRA at a per-run budget of ε=8 and compare, on the same prompts and at the same budget, a control that removes only the reward signal and two private supervised fine-tuning recipes. The reward signal improves accuracy over the control by 2.65 points on MATH and 3.24 on GSM8K, in every seed; the improvement survives a format-robust scorer, at 1.3 points on MATH, and is not explained by response length. At the same budget the private model outperforms both supervised recipes on MATH and GSM8K by 2.3 to 3.8 points, retains 85--90% of the gain of non-private GRPO on these tasks, and on MATH the noise of an eightfold tighter budget costs at most 1.2 points. The reward effect also carries to CommonsenseQA, an exploratory non-mathematical task. Verifier feedback thus remains a usable learning signal under prompt-level privacy.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. An Imperfect Verifier is Good Enough: Learning with Noisy Rewards

    Apr 9, 2026Andreas Plesner, Francisco Guzmán, Anish AthalyeRL for Code GenerationReinforcement Learning

  2. Quantifying Empirical Compute-Supervision Tradeoffs in RLVR

    May 24, 2026Ryo Mitsuhashi, Patrick Chen, Isabelle Tseng +2Reinforcement LearningRL for Language Model Reasoning