Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.
Figures & tables
Figure 1: The separation is present; the flatness is not. Training reward for the published easy, learnable, and unlearnable cohorts. Slopes per 100 optimizer steps are +0.3800 , +0.1119 , and +0.0422[+0.0259,+0.0590] , respectively. The affected cohort remains markedly slower to improve, but its slope is not zero.
Figure 2: The set is unstable in both seed count and aggregation operator. Top: set size changes as more seeds are combined; many prompts are flagged by only a subset of seeds, with disagreement concentrated near τ=0.1 . Bottom: holding the same 292 sub-threshold candidates fixed, the published five-seed construction returns 74 prompts, individual seeds return 145–150 (mean 147.8), a zero-correct-in-128 rule returns 181, and the alternative five-seed aggregation returns 228.
Figure 3: The sampling budget can be predicted before it is spent. (a) Predicted Jaccard agreement between two independent threshold-defined sets as a function of per-prompt budget N ; the lower panel shows the discrete effective threshold that produces the small- N sawtooth. (b) Predicted versus measured Jaccard. The filled diamonds are the two prospectively frozen N=40 and N=64 predictions. Retrospective and cross-model consistency checks are shown for context; the explicitly excluded 1.5B point is not treated as a validation result.
Figure 4: Gradient similarity after matching estimator sample count. Left: within-group mean cosine as the number K of correct rollouts averaged per prompt is matched across cohorts. The fixed matched cohort contains 28 Du , 28 learnable, and 38 easy prompts. Right: the easy-to- Du similarity ratio. Matching reduces the unmatched 2.327 times ratio to approximately 1.5 times, while every tested matched level remains above one. Open markers show each group’s natural, unmatched sample count. At K=11 , matching retains only 28 of the 71 Du prompts in the gradient cohort; the matched values therefore demonstrate sensitivity to unequal estimator sample count in an eligible subset, rather than providing corrected full-cohort estimates.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Control
Comparison
Result
Interpretation
Same-model resampling
fixed checkpoint vs. cross-seed, N=32
0.798 vs. 0.751
Most apparent cross-seed disagreement can arise from evaluation sampling alone.
Token-cap sensitivity
1,024 vs. 5,120 tokens
0.898 Jaccard
Set movement is no larger than the measured repeatability scale.
Equal-cost allocation
one N=128 pass vs. 4×32
0.877 vs. 0.791
A deeper estimate is more reproducible at the same rollout cost.
Appendix
Table 1: Key robustness controls. The first control separates evaluation noise from training variation; the second tests the generation-cap confound; the third asks how an equal sampling budget should be allocated.
Figure 5: Equal-budget comparison of one deep N=128 evaluation and four shallow blocks of 32. The deeper rule is more reproducible both before and after applying the never-rewarded exclusion. Error bars show 95% intervals.
Model / rule
Measured
Predicted
Error
Qwen2.5-0.5B, one pass N=32
0.7977
0.8000
+0.0023
Qwen2.5-0.5B, one pass N=128
0.8770
0.8759
-0.0011
Qwen2.5-0.5B, four blocks of 32
0.7906
0.8228
+0.0322
Qwen2.5-0.5B, two blocks of 16
0.7028
0.7248
+0.0220
Llama-3.2-3B, two blocks of 16
0.6844
0.6763
-0.0081
Llama-3.2-3B v2, two blocks of 16
0.7070
0.6926
-0.0144
Appendix
Table 2: Retrospective and cross-model consistency checks for the sampling-budget model. Six rows are treated as retrospective or consistency checks; the Qwen2.5-1.5B row is an excluded diagnostic and is not counted as validation evidence. The cross-model rows reuse the generation pass from which each pass-rate distribution was estimated and are therefore consistency checks rather than independent validation results.
Figure 6: Predicted versus measured Jaccard agreement across retrospective checks, cross-model consistency checks, and prospectively frozen predictions. Filled diamonds mark the two prospective measurements; the Qwen2.5-1.5B point is shown separately and is not counted as validation evidence.
Figure 7: Full 95% interval width of the easy-to- Du gradient-similarity ratio induced by redrawing the threshold-defined set from the validated sampling model, as a function of per-prompt budget N . The Du -anchored variant is primary and the label-faithful variant secondary. The dotted line marks the pre-specified 0.25 materiality threshold; the dashed line marks the 0.50 sampling confidence-interval width already reported for the downstream statistic. At N=32 the two anchorings straddle the threshold.
Configuration constant
Ours
Original
Training
Dataset name
simplelr_qwen_level1to4_sub1k
simplelr_qwen_level1to4
Maximum response length
1,024
5,120
Train batch size
32
256
PPO mini-batch size
16
64
PPO micro-batch size
2
32
Appendix
Table 3: Configuration differences identified from the executed training and evaluation paths. The main text highlights differences that directly affect scientific interpretation; this table records the complete audit.
Reinforcement Learning with Verifiable Reward (RLVR) has proven effective in improving Large Language Model's (LLM) reasoning ability. However, the learning dynamics of RLVR remain underexplored. In this paper, we reveal a counterintuitive phenomenon: among hard examples that the model initially struggles with, a substantial subset remains unlearnable even when correct rollouts are present. To understand the phenomenon, we first demonstrate that existing optimization and sampling techniques fail to resolve unlearnability. With cross-example gradient analysis, we show that unlearnable examples have fundamental representation issue, characterized by low gradient similarity with the rest of the examples and ungeneralizable reasoning patterns. We further show that representation flaws are difficult to mitigate in RL, as data augmentation does not improve gradient similarity. Our study provides the first systematic characterization of unlearnable data in RLVR training and reveals fundamental limitations in current RL approaches for reasoning tasks. Code and data are available at \url{https://github.com/yulinchen99/unlearnability-rlvr}.
Reinforcement Learning with Verifiable Reward (RLVR) is empirically shown to notably enhance the reasoning performance of large language models (LLMs), particularly in mathematics and programming. However, the mechanistic role of Sample Difficulty in RLVR remains poorly understood. In this paper, we investigate RLVR through the lens of difficulty-wise and one-sample analysis. We find that sample difficulty has a non-monotonic effect on RLVR: easy and medium-difficulty problems yield the strongest and most stable reasoning improvements, whereas overly hard problems often provide weak learning signals, induce degenerate behaviors such as answer repetition or skipping necessary computation, and can ultimately degrade the model's pre-existing capabilities. Beyond the obverse of response, we further analyze the model's internal feature dynamics using Temporal Sparse Autoencoders (T-SAE). Easy problems mainly reinforce direct-answer and basic-computation features while suppressing deliberative-reasoning features; hard problems activate reasoning-related features but become useful only when successful trajectories are sampled; medium-difficulty problems provide a more balanced signal, strengthening both computation and multi-step reasoning features. Motivated by these findings, we propose difficulty-adaptive strategies for hard-sample utilization, using backward-reasoning reformulation and T-SAE-guided training signals to improve reward density and credit assignment during RLVR. Overall, our results identify sample difficulty as a key factor governing both the optimization dynamics and representation evolution of RLVR.
Yue Cheng, Jiajun Zhang, Xiaohui Gao +3
Beijing Jiaotong University · AntGroup · Northwestern Polytechnical University +2
Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it remains a mystery how rewards based solely on final outcomes can help overcome the long-horizon barrier to extended reasoning. To understand this, we develop a theory of the training dynamics of RLVR for transformers on compositional reasoning tasks. Our theory shows that mixed-difficulty training naturally induces an implicit curriculum: without any explicit schedule, easier problems become learnable first and shape the frontier for harder ones, creating a learning progression from easy to hard during optimization. The effectiveness of this curriculum is governed by the smoothness of the difficulty spectrum. When the spectrum is smooth, training dynamics enter a well-behaved relay regime, in which persistent gradient signals on easier problems make slightly harder ones tractable and keep training at the edge of competence. When the spectrum contains abrupt discontinuities, training undergoes grokking-type phase transitions with prolonged plateaus before progress recurs. As a technical contribution, our analysis develops and adapts techniques from Fourier analysis on finite groups to our setting. We validate the predicted mechanisms empirically via controlled synthetic experiments and real-model RLVR runs.
Yu Huang, Zixin Wen, Yuejie Chi +4
Department of Statistics and Data Science, Wharton School, University of Pennsylvania · Machine Learning Department, Carnegie Mellon University · Department of Statistics and Data Science, Yale University +1