Frontier Learning: Training LLM Reasoners at the Edge of Capability
Organizations: University College London London, UK · University of Basel Basel, Switzerland
Abstract
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.
Figures & tables
| Puzzle | Math | ||
| Task | Countdown | Sokoban | Dec. Arith. |
| Model | Qwen3-4B-Base | Qwen3-4B | Qwen3-4B-Base |
| Method | |||
| Domain Rand. | 44.7 ±1.7 | 37.6 ±2.6 | 31.0 ±2.8 |
| Uniform | 44.2 ±4.0 | 39.5 ±1.3 | 28.0 ±2.6 |
| SEC | 40.1 ±8.1 | 34.3 ±6.6 | 29.1 ±4.1 |
| Method | Accuracy |
|---|---|
| Domain Rand. | 23.4 ±4.7 |
| Uniform | 31.9 ±1.5 |
| SEC | 33.3 ±1.2 |
| PLR | 31.2 ±1.9 |
| ACE-GRPO | 33.2 ±2.1 |
| DAPO | 25.4 ±5.9 |
| Task | Countdown | Largest Island |
|---|---|---|
| Model | Llama-3.2-3B-Instruct | Olmo3-7B-Instruct |
| Method | ||
| Domain Rand. | 33.8 ±1.9 | 45.2 ±1.8 |
| Uniform | 25.3 ±4.7 | 45.3 ±3.4 |
| SEC | 29.5 ±1.6 | 48.3 ±4.7 |
| PLR | 29.0 ±1.8 | 52.0 ±5.3 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Level | Bin | Config | ||
|---|---|---|---|---|---|
| Countdown | 1 | Easy | num =3 | mv =20 | mt =80 |
| 2 | Easy | 3 | 40 | 150 | |
| 3 | Easy | 3–4 | 60 | 250 | |
| 4 | Medium | 4 | 80 | 350 | |
| 5 | Medium | 4–5 | 100 | 500 | |
| 6 | Medium | 5 | 120 | 800 | |
| Hyperparameter | Symbol | Value |
|---|---|---|
| Learning rate | ||
| Train batch size | 64 | |
| Rollouts per problem | 8 | |
| PPO mini-batch size | 16 | |
| PPO micro-batch size per GPU | 8 | |
| Max problem length (tokens) | 1024 |
| Hyperparameter | Symbol | Value |
|---|---|---|
| Buffer and sampling | ||
| Levels per step | 4 | |
| Problems per level | 16 | |
| Buffer capacity | 100 | |
| Seed levels | 8 | |
| Seed sampling mode | grid | |
| Variant | Accuracy (%) |
|---|---|
| Default frontier learning | |
| Informative interval | |
| Informative interval | |
| Method | Runtime (h) | Response length (tokens) | Accuracy (%) |
|---|---|---|---|
| Domain Rand. | 15.4 2.0 | 3838 33 | 37.6 2.6 |
| PLR | 14.8 1.8 | 3698 103 | 41.1 3.9 |
| Uniform | 15.1 1.8 | 3761 51 | 39.5 1.3 |
| SEC | 14.8 1.8 | 3709 81 | 34.3 6.6 |
| Frontier Learning (ours) | 14.2 1.7 | 3529 49 | 45.2 1.9 |
| Method | Runtime (h) | Response length (tokens) | Accuracy (%) |
|---|---|---|---|
| Domain Rand. | 17.4 0.1 | 1538 14 | 44.7 1.7 |
| PLR | 11.3 0.3 | 1521 44 | 44.3 2.9 |
| Uniform | 11.8 0.3 | 1546 50 | 44.2 4.0 |
| SEC | 11.2 0.4 | 1446 65 | 40.1 8.1 |
| Frontier Learning (ours) | 11.7 0.5 | 1269 28 | 50.9 0.5 |
| Method | Runtime (h) | Response length (tokens) | Accuracy (%) |
|---|---|---|---|
| Domain Rand. | 4.3 1.1 | 837 31 | 31.0 2.8 |
| PLR | 6.0 1.1 | 677 30 | 27.5 2.9 |
| Uniform | 5.7 1.0 | 642 33 | 28.0 2.6 |
| SEC | 6.1 1.2 | 665 54 | 29.1 4.1 |
| Frontier Learning (ours) | 4.2 1.1 | 758 50 | 34.7 3.6 |
| Method | Runtime (h) | Response length (tokens) | Accuracy (%) |
|---|---|---|---|
| Baselines | |||
| Domain Rand. | 19.4 1.9 | 907 104 | 23.4 4.7 |
| PLR | 18.7 1.1 | 877 65 | 31.2 1.9 |
| Uniform | 20.8 1.9 | 994 108 | 31.9 1.5 |
| SEC | 18.9 2.3 | 866 129 | 33.3 1.2 |
| ACE-GRPO | 22.0 0.8 | 1063 45 | 33.2 2.1 |
| Method | Accuracy (%) |
|---|---|
| DR | 25.4 3.7 |
| PLR | 31.2 1.9 |
| Uniform | 32.2 1.7 |
| SEC | 33.5 1.6 |
| AceGRPO | 31.8 3.3 |
| Frontier Learning (ours) | 68.0 5.5 |
| Method | Runtime (h) | Response length (tokens) | Accuracy (%) |
|---|---|---|---|
| Domain Rand. | 63.2 0.3 | 1578 29 | 45.2 1.8 |
| PLR | 44.3 6.0 | 1276 96 | 52.0 5.3 |
| Uniform | 53.1 2.5 | 1132 125 | 45.3 3.4 |
| SEC | 44.4 7.1 | 1218 102 | 48.3 4.7 |
| ACE-GRPO | 43.8 4.7 | 1243 106 | 45.3 5.5 |
| frontier learning (ours) | 52.4 0.7 | 1498 6 | 73.7 3.3 |