cs.LGAug 18, 2026

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

Authors: Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez

Organizations: University of Valencia

Abstract

Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and a rubric reward that selects broad-topic answering with low semantic leakage under held-out evaluation.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

    Jul 30, 2026Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj +1Catastrophic ForgettingComparison

  2. Replay What Matters: Off-Policy Replay for Efficient LLM Reinforcement Unlearning

    Jun 13, 2026Zirui Pang, Chenlong Zhang, Haosheng Tan +3Large Language Model UnlearningLarge Language Model Reinforcement Learning

  3. Null-Space Constrained Low-Rank Adaptation for Response-Specified Large Language Model Unlearning

    Jun 9, 2026Bocheng Ju, Jianhua Wang, Chengliang Liu +1Large Language Model UnlearningExact Unlearning