Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR
Organizations: Department of Robotics and Mechatronics Engineering, University of Dhaka
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.
Figures & tables
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Control | Comparison | Result | Interpretation |
|---|---|---|---|
| Same-model resampling | fixed checkpoint vs. cross-seed, | 0.798 vs. 0.751 | Most apparent cross-seed disagreement can arise from evaluation sampling alone. |
| Token-cap sensitivity | 1,024 vs. 5,120 tokens | 0.898 Jaccard | Set movement is no larger than the measured repeatability scale. |
| Equal-cost allocation | one pass vs. | 0.877 vs. 0.791 | A deeper estimate is more reproducible at the same rollout cost. |
| Model / rule | Measured | Predicted | Error |
|---|---|---|---|
| Qwen2.5-0.5B, one pass | 0.7977 | 0.8000 | +0.0023 |
| Qwen2.5-0.5B, one pass | 0.8770 | 0.8759 | -0.0011 |
| Qwen2.5-0.5B, four blocks of 32 | 0.7906 | 0.8228 | +0.0322 |
| Qwen2.5-0.5B, two blocks of 16 | 0.7028 | 0.7248 | +0.0220 |
| Llama-3.2-3B, two blocks of 16 | 0.6844 | 0.6763 | -0.0081 |
| Llama-3.2-3B v2, two blocks of 16 | 0.7070 | 0.6926 | -0.0144 |
| Configuration constant | Ours | Original |
|---|---|---|
| Training | ||
| Dataset name | simplelr_qwen_level1to4_sub1k | simplelr_qwen_level1to4 |
| Maximum response length | 1,024 | 5,120 |
| Train batch size | 32 | 256 |
| PPO mini-batch size | 16 | 64 |
| PPO micro-batch size | 2 | 32 |