Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.
Figures & tables
Figure 1: Frontier learning with procedural data generation. Levels vary along task-specific difficulty dimensions, here abstracted as structural complexity and reasoning depth. As the model improves, its frontier of capabilities shifts outward; frontier learning uses procedural generation to continually construct training data near this moving region.
Figure 2: One iteration of frontier learning with GRPO. At each step, the algorithm initially constructs the level batch with a mix of exploratory levels and high-regret frontier levels sampled from the buffer. A level sampled from the buffer is locally mutated with probability pmut(ℓ) , by perturbing its attributes in order to evaluate and potentially extend the policy’s capability frontier. The resulting level batch is then passed through the generator to construct the training batch of problems for the GRPO update. The resulting rewards and regret estimates are then used for future level selection.
Puzzle
Math
Task
Countdown
Sokoban
Dec. Arith.
Model
Qwen3-4B-Base
Qwen3-4B
Qwen3-4B-Base
Method
Domain Rand.
44.7 ±1.7
37.6 ±2.6
31.0 ±2.8
Uniform
44.2 ±4.0
39.5 ±1.3
28.0 ±2.6
SEC
40.1 ±8.1
34.3 ±6.6
29.1 ±4.1
Table 1: Accuracy (%) on the fixed anchor evaluation set after 200 training steps. Best result per task in bold ; second-best underlined .
Method
Accuracy
Domain Rand.
23.4 ±4.7
Uniform
31.9 ±1.5
SEC
33.3 ±1.2
PLR
31.2 ±1.9
ACE-GRPO
33.2 ±2.1
DAPO
25.4 ±5.9
Table 2: Accuracy (%) on Dice for the baselines, Frontier Learning and its components ablation. Best result in bold ; second-best within each block underlined .
Task
Countdown
Largest Island
Model
Llama-3.2-3B-Instruct
Olmo3-7B-Instruct
Method
Domain Rand.
33.8 ±1.9
45.2 ±1.8
Uniform
25.3 ±4.7
45.3 ±3.4
SEC
29.5 ±1.6
48.3 ±4.7
PLR
29.0 ±1.8
52.0 ±5.3
Table 3: Accuracy (%) on Countdown and Largest Island . Best result per task in bold ; second-best underlined .
Figure 3: Per-difficulty accuracy (%) on Largest Island with Olmo3-7B-Instruct over 500 GRPO steps. All methods saturate on Easy levels, but only frontier learning keeps improving on Medium and Hard levels throughout training, while the baselines plateau after roughly 200 steps.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Example Sokoban levels at the two extremes of the anchor ladder. Left: L1 — grid 6×6 , 1 box, depth ≤12 . Right: L7 — grid 10×10 , 6 boxes, depth ≤80 . Colours: dark grey = wall, beige = floor, brown = box, yellow = goal, green = box on goal.
Task
Level
Bin
Config
Countdown
1
Easy
num =3
mv =20
mt =80
2
Easy
3
40
150
3
Easy
3–4
60
250
4
Medium
4
80
350
5
Medium
4–5
100
500
6
Medium
5
120
800
Appendix
Table 4: Full anchor-level specifications, grouped into Easy/Medium/Hard (and Extra Hard for Dice) difficulty bins used throughout the paper. Countdown: num = num_numbers , mv = max_value , mt = max_target . Sokoban: grid is square; d = max_depth . Decimal: dp = decimal_places , pr = precision , t = terms . Largest Island: grid = rows × cols ; isl = num_islands ; sz = island_size .
Hyperparameter
Symbol
Value
Learning rate
10−6
Train batch size
64
Rollouts per problem
nr
8
PPO mini-batch size
16
PPO micro-batch size per GPU
8
Max problem length (tokens)
1024
Appendix
Table 5: GRPO training hyperparameters (shared across all methods and tasks).
Figure 5: Accuracy (%) on the fixed anchor evaluation sets over 200 GRPO steps for Countdown , Sokoban , and Decimal Arithmetic .
Figure 6: Validation accuracy across the three difficulty bins throughout the full training horizon for Countdown , Sokoban , and Decimal Arithmetic .
Figure 7: Final validation accuracy per difficulty bin for Countdown , Sokoban , and Decimal Arithmetic .
Figure 8: Long-horizon Dice learning curves over 500 GRPO steps. DAPO is omitted from the step-based curves because it is matched by total rollout budget rather than optimizer updates.
Figure 9: Validation accuracy across the four difficulty bins throughout the full training horizon for Dice .
Figure 10: Final validation accuracy per difficulty bin for Dice
Variant
Accuracy (%)
Default frontier learning
39.2±1.3
λs=0.02
40.7±1.1
Informative interval =[0.10,0.90]
40.5±0.8
Informative interval =[0.02,0.98]
39.8±1.0
λs=0
38.8±1.4
Appendix
Table 7: Hyperparameter sensitivity on Countdown using Llama-3.2-3B-Instruct.
Figure 11: Learning-rate sweep on Dice (seed 42) for Uniform and SEC, the non-regret-based baseline and the best-performing baseline in Table 2 , respectively. † Denotes the learning rate used for the main paper’s results ( 1×10−6 ). No swept value raises either baseline’s final accuracy above ≈ 35%, indicating that an under-tuned learning rate does not explain the gap to frontier learning.
Method
Runtime (h)
Response length (tokens)
Accuracy (%)
Domain Rand.
15.4 ± 2.0
3838 ± 33
37.6 ± 2.6
PLR
14.8 ± 1.8
3698 ± 103
41.1 ± 3.9
Uniform
15.1 ± 1.8
3761 ± 51
39.5 ± 1.3
SEC
14.8 ± 1.8
3709 ± 81
34.3 ± 6.6
Frontier Learning (ours)
14.2 ± 1.7
3529 ± 49
45.2 ± 1.9
Appendix
Table 8: End-to-end compute accounting on Sokoban (200 steps).
Method
Runtime (h)
Response length (tokens)
Accuracy (%)
Domain Rand.
17.4 ± 0.1
1538 ± 14
44.7 ± 1.7
PLR
11.3 ± 0.3
1521 ± 44
44.3 ± 2.9
Uniform
11.8 ± 0.3
1546 ± 50
44.2 ± 4.0
SEC
11.2 ± 0.4
1446 ± 65
40.1 ± 8.1
Frontier Learning (ours)
11.7 ± 0.5
1269 ± 28
50.9 ± 0.5
Appendix
Table 9: End-to-end compute accounting on Countdown (200 steps).
Method
Runtime (h)
Response length (tokens)
Accuracy (%)
Domain Rand.
4.3 ± 1.1
837 ± 31
31.0 ± 2.8
PLR
6.0 ± 1.1
677 ± 30
27.5 ± 2.9
Uniform
5.7 ± 1.0
642 ± 33
28.0 ± 2.6
SEC
6.1 ± 1.2
665 ± 54
29.1 ± 4.1
Frontier Learning (ours)
4.2 ± 1.1
758 ± 50
34.7 ± 3.6
Appendix
Table 10: End-to-end compute accounting on Decimal Arithmetic (200 steps).
Method
Runtime (h)
Response length (tokens)
Accuracy (%)
Baselines
Domain Rand.
19.4 ± 1.9
907 ± 104
23.4 ± 4.7
PLR
18.7 ± 1.1
877 ± 65
31.2 ± 1.9
Uniform
20.8 ± 1.9
994 ± 108
31.9 ± 1.5
SEC
18.9 ± 2.3
866 ± 129
33.3 ± 1.2
ACE-GRPO
22.0 ± 0.8
1063 ± 45
33.2 ± 2.1
Appendix
Table 11: End-to-end compute accounting on Dice (500 steps).
Method
Accuracy (%)
DR
25.4 ± 3.7
PLR
31.2 ± 1.9
Uniform
32.2 ± 1.7
SEC
33.5 ± 1.6
AceGRPO
31.8 ± 3.3
Frontier Learning (ours)
68.0 ± 5.5
Appendix
Table 12: Dice performance at matched PLR compute time (18.7 hours).
Method
Runtime (h)
Response length (tokens)
Accuracy (%)
Domain Rand.
63.2 ± 0.3
1578 ± 29
45.2 ± 1.8
PLR
44.3 ± 6.0
1276 ± 96
52.0 ± 5.3
Uniform
53.1 ± 2.5
1132 ± 125
45.3 ± 3.4
SEC
44.4 ± 7.1
1218 ± 102
48.3 ± 4.7
ACE-GRPO
43.8 ± 4.7
1243 ± 106
45.3 ± 5.5
frontier learning (ours)
52.4 ± 0.7
1498 ± 6
73.7 ± 3.3
Appendix
Table 13: End-to-end compute accounting on Largest Island (500 steps, Olmo-3-7B-Instruct).
Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not teach new strategies; it redistributes probability mass over solutions the base model already contains. In this work, we ask: if RL merely steers the model toward paths it already knows, is the RL optimization loop itself necessary? Through token-level analysis across multiple model families and RL algorithms, we find that RL's beneficial footprint is a sparse, predictable correction concentrated at high-entropy decision points where the model is uncertain which branch to take. Only 1--3% of token positions are affected, the promoted token always lies within the base model's top-5 alternatives, and targeted corrections at those few positions causally recover a large fraction of RL's accuracy gain, while random corrections fail. The base model's own entropy identifies these positions without any RL-trained model, and the entire correction is low-dimensional, representable in a tiny fraction of model parameters. These findings reframe reasoning improvement as sparse policy selection, not capability acquisition. We translate this insight into ReasonMaxxer, a minimal RL-free method that applies contrastive loss only at entropy-gated decision points, using a few hundred base-model rollouts and no online generation. Across three model families, six scales, and six math reasoning benchmarks, ReasonMaxxer matches or exceeds full RL performance while requiring only tens of problems and minutes of single-GPU training, a reduction in training cost of roughly three orders of magnitude.
Ömer Faruk Akgül, Rajgopal Kannan, Willie Neiswanger +1
Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks. Recent research suggests that Chain-of-Thought (CoT) reasoning paths are inherent in pre-trained LLMs and can be elicited by simply altering the decoding process, where the presence of a CoT path correlates with higher answer confidence. Building on these insights, we present Reinforcement Learning from Self-Feedback (RLSF), a post-training stage that utilises the model's intrinsic confidence as a self-generated reward. By generating multiple CoT decoding beams from a frozen LLM, we compute the confidence of each final answer span and rank the resulting traces accordingly to create synthetic preferences. These preferences are subsequently utilised to fine-tune the policy through standard preference optimisation, requiring no human labels, gold answers, or externally curated rewards. RLSF simultaneously (i) refines the model's probability estimates--restoring well-behaved calibration--and (ii) strengthens step-by-step reasoning, yielding improved performance on arithmetic reasoning and multiple-choice question answering. By converting a model's own uncertainty into structured self-feedback, RLSF affirms reinforcement learning on intrinsic model behaviour as a principled and data-efficient component of the LLM post-training pipeline. Our results demonstrate that leveraging these inherent reasoning capabilities provides a robust path for enhancing model reliability without manual prompt engineering or external supervision.
Carel van Niekerk, Renato Vukovic, Benjamin Ruppik +3
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across both stages are prohibitively expensive. To address these challenges, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline. We follow the standard LLM training pipeline by pretraining language models from 5M to 1B parameters on human chess games, supervised fine-tuning on synthetic reasoning traces, and running RL on chess puzzles with verifiable rewards. Using this framework, we find that the post-RL performance at given RL compute level is well-predicted from the pretraining loss, and slope of the RL reward curves improves approximately linearly with the pretraining tokens. Beyond scaling, we find that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves the SFT policy already preferred, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. We further test whether our findings transfer beyond chess by training a 1B language model on math-domain text, where the same predictive pattern emerges: longer-pretrained checkpoints reach higher post-RL performance and improve faster under RL. In sum, we provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying the science of reasoning across the full pretraining-to-post-training pipeline.
Jingyan Shen, Ang Li, Salman Rahman +4
New York University · Columbia University · University of California, Los Angeles +2