cs.LGSep 28, 2026

Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training

Authors: Yuyang Deng, Yu Wang, Jiayun Wang

Organizations: Accenture, Center for Advanced AI · Georgia Institute of Technology

Abstract

Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}volving \textbf{O}ptimization (DEO), which replaces challenger parameter updates with solver-guided task sampling. The KL-regularized challenger objective defines an exponential tilt of a fixed base task distribution. DEO uses this distribution as a sampling target: a frozen LLM generates and mutates tasks, the solver scores them, and an approximate Metropolis selection rule refines the training pool. Only the solver is trained. Theoretically, for an idealized variant that samples exactly from the tilted distribution, and under regularity, local gradient-dominance, and initialization conditions, we show that DEO learns distributionally robust reasoning ability. In experiments, DEO achieves reasoning performance competitive with R-Zero while using over 50%50\% less wall-clock training time, and improves reasoning accuracy over a no-walk ablation. Replacing the task generator with a frozen API-only LLM further improves the local solver, illustrating a capability enabled by removing challenger training.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. INFUSER: Influence-Guided Self-Evolution Improves Reasoning

    Jun 8, 2026Siyu Chen, Miao Lu, Beining Wu +7Self-EvolutionIterative Co-Training

  2. A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

    Oct 21, 2025Mengqi Li, Lei Zhao, Anthony Man-Cho So +2LLM Reasoning StrategiesMathematical Reasoning Benchmarks

  3. Learning What to Practice: Diagnosis-Guided Self-Evolution for Language Models

    Sep 1, 2026Xincheng Wei, Yifan Ding, Fucheng Xiong +6Self-EvolutionMathematical Reasoning Benchmarks