cs.LGSep 28, 2026

Frontier Learning: Training LLM Reasoners at the Edge of Capability

Authors: Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi

Organizations: University College London London, UK · University of Basel Basel, Switzerland

Abstract

Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

    May 7, 2026Ömer Faruk Akgül, Rajgopal Kannan, Willie Neiswanger +1LLM Reasoning StrategiesReasoning Skills

  2. Post-Training Large Language Models via Reinforcement Learning from Self-Feedback

    Jul 29, 2025Carel van Niekerk, Renato Vukovic, Benjamin Ruppik +3Large Language Model Reinforcement LearningLLM Reasoning Strategies

  3. Understanding Reasoning from Pretraining to Post-Training

    Jul 17, 2026Jingyan Shen, Ang Li, Salman Rahman +4LLM Reasoning StrategiesLarge Language Model Pretraining