What pass@k Cannot Measure: Evaluating Diversity and Capability Retention after Post-Training
Organizations: Accenture Strategy & Consulting · Vizuara AI Labs (vizuara.com) · Vizuara AI Labs (vizuara.ai)
Abstract
pass@, the fraction of problems a model solves within sampled attempts, is the field's default protocol for deciding whether reinforcement-learning (RL) post-training on verifiable rewards improved a model. At the population level, pass@ depends only on a problem's probability of a correct sample, with no term for how it is distributed across outputs. We show this gap is not academic. Training Qwen2.5-1.5B-Instruct on grade-school math with Group Relative Policy Optimization (GRPO) and with rejection-sampling fine-tuning (RFT, training on the model's own shortest verifier-passed rollout) moves three complementary diversity measures (token-level entropy, answer-level entropy, unique answers per prompt) in opposite directions, with zero overlap across three seeds per arm. The gap survives restricting to verifier-correct completions only (lexical diversity among correct solutions is 15% lower for GRPO, after controlling for length) and a count-controlled check isolating diversity among incorrect answers alone, ruling out that GRPO's higher accuracy alone explains it. Yet pass@8 and pass@32 show no consistent winner on GSM8K, and a hard MATH-500 subset shows the same pattern: separation only at low . Compared against the starting checkpoint, no trained arm significantly improves hard-problem coverage: RFT is significantly worse, while GRPO is statistically indistinguishable from it - so GRPO's pass@1 edge over RFT reflects a smaller loss relative to Base, not a capability gain, a missing-control issue, not a failure of pass@. On GSM8K, only pass@1, with no role in detecting diversity by construction, separates the arms cleanly, rewarding the arm whose correct solutions are least diverse. We argue this is a concrete instance of a standard evaluation protocol missing a property it is routinely used to certify.
Figures & tables
| Arm | Token entropy | Unique answers/problem | Answer entropy |
|---|---|---|---|
| RFT (shortest-correct) | |||
| RFT + KL brake | |||
| GRPO | |||
| Single seed each; context only, not a robustness claim: | |||
| GRPO + length penalty ( ) | |||
| GRPO + length penalty ( ) | |||
| Arm | pass@1 | pass@8 | pass@32 |
|---|---|---|---|
| RFT (shortest-correct) | [ , ] | [ , ] | [ , ] |
| RFT + KL brake | [ , ] | [ , ] | [ , ] |
| GRPO | [ , ] | [ , ] | [ , ] |
| GRPO + length penalty ( ) | [ , ] | [ , ] | [ , ] |
| GRPO + length penalty ( ) | [ , ] | [ , ] | [ , ] |
| OPD | [ , ] | [ , ] | [ , ] |