Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-generated questions: invalid questions and repeated variants of the same mathematical questions. First, invalid questions become more prevalent across rounds, and answer-consistency filtering further increases their proportion in training data. Second, existing question diversity controls based on lexical similarity can miss mathematically equivalent questions expressed in different ways, which leads to question diversity collapse in later training rounds. Building on these findings, we introduce R-Quest, which uses question validity and novelty feedback to guide self-evolution. We first train the solver to recognize and reject invalid questions, then use its judgments to guide questioner rewards and filter solver training data. To avoid question repetition, we use a frozen base model to compare sampled question pairs and provide novelty feedback. Empirically, our method consistently achieves the highest average performance on 12 benchmarks in mathematical reasoning, general-domain reasoning, and code generation across two model families. Additionally, R-Quest maintains stable performance gains over ten rounds of self-evolution, peaking in the final round and outperforming R-Zero by 17.32 points.
Figures & tables
Figure 1: Average performance on seven math benchmarks. R-Quest exceeds the peak of R-Zero in every round and peaks at round ten, leading by 17.32 points.
Figure 2: Two question-quality problems in self-evolution and how R-Quest addresses them: (a) R-Quest can recognize and reject invalid questions; (b) R-Quest can detect repeated mathematical tasks despite differences in expression.
Figure 3: Question validity rate and solver performance on Qwen3-4B-Base. (a) Valid-question rates before and after R-Zero’s answer-consistency filter. (b) Average math reasoning performance of the solver trained on matched original and repaired datasets.
Figure 4: Repeated question distribution and BLEU fragmentation in the fifth round of R-Zero. BLEU separates the illustrated mathematically equivalent questions into multiple different clusters.
Method
Avg.
AMC
Minerva
MATH
GSM8K
Olympiad
AIME25
AIME24
Qwen3-4B-Base
Base model
46.17
47.34
51.84
74.20
88.63
39.11
11.35
10.73
Base model (validity init.)
47.07
49.14
53.68
74.40
89.76
41.04
10.73
10.73
R-Zero
49.10
54.84
54.78
76.20
92.42
42.37
12.29
10.83
Table 1: Mathematical reasoning results (%) across seven benchmarks. Avg. denotes the mean across benchmarks. For each backbone, the best score in each column is shown in bold .
Figure 5: Average performance on seven benchmarks over ten rounds on Qwen3-4B-Base. Yellow and orange denote R-Quest without novelty or validity feedback.
Method
Code generation
General-domain reasoning
Avg.
HumanEval+
MBPP+
Avg.
SuperGPQA
MMLU-Pro
BBEH
Qwen3-4B-Base
Base model
58.79
54.88
62.70
28.79
26.21
51.41
8.74
Base model (validity init.)
59.63
55.49
63.76
29.40
26.28
51.58
10.35
Table 2: Code generation (pass@1) and general-domain reasoning results (%). Avg. denotes the mean within each domain. Best scores within each backbone are in bold .
Variant
Math
Code
General
R-Quest
51.64
62.69
32.75
w/o novelty
51.02
61.98
31.80
w/o validity
50.27
61.28
31.53
w/o both (R-Zero)
49.10
59.27
31.21
Table 3: Component ablations on Qwen3-4B-Base. Scores are averaged within each domain.
Figure 6: Validity assessment, question and answer quality during self-evolution on Qwen3-4B-Base. Panels (a) and (b) show R-Quest’s improved validity assessment on held-out questions. Panels (c) and (d) show that R-Quest maintains high question validity and better preserves majority-answer accuracy on valid questions. Bands in (c) show pointwise 95% Wilson confidence intervals.
Figure 7: Question type clusters (bars, left axis) and mathematical performance (lines, right axis) on Qwen3-4B-Base. Numbers above the bars denote top-five group proportion.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Error type
Count
Percentage
Contradiction
380
28.9%
No solution
362
27.5%
Ambiguity
315
23.9%
Missing information
181
13.7%
Poorly formed questions
72
5.5%
Other
7
0.5%
Appendix
Table 5: Error types among the 1,317 invalid questions in the training and validation splits.
Setting
Value
Training steps
15
Questions per rollout batch
512
Responses per question
8
Actor global batch size
128
Learning rate
10−6
KL coefficient
10−2
Appendix
Table 6: Training settings for validity-aware solver initialization.
Figure 8: Approximate pass probability (1−p)K for the default comparison budget K=8 and the additional setting K=16 .
Table 14
Variant
Mathematical reasoning
Avg.
AMC
Minerva
MATH
GSM8K
Olympiad
AIME25
AIME24
R-Quest
51.64
57.89
58.82
77.60
92.57
46.67
13.44
14.48
w/o novelty
51.02
55.94
56.62
78.00
91.81
45.63
15.42
13.75
w/o validity
50.27
56.48
58.82
77.00
91.74
41.19
13.54
13.12
w/o both (R-Zero)
49.10
54.84
54.78
76.20
92.42
42.37
12.29
10.83
Appendix
Table 9: Per-benchmark results for the gate ablations on Qwen3-4B-Base.
Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce INFUSER, an iterative co-training framework with two co-evolving roles: a Generator that drafts questions and reference golden answers from a pool of unstructured, automatically collected documents, and a Solver that improves by training on them. The solver is trained with standard correctness rewards against the generator-provided answers, while the generator is rewarded by an optimizer-aware influence score that measures whether each proposed question would actually improve the solver on the target distribution. Because this continuous, noisy influence score is poorly served by standard GRPO, we propose DuGRPO, a dual-normalized variant of GRPO, for generator training. Together, these turn the document pool into an adaptive curriculum that favors questions useful to the current solver, not just hard ones. On Qwen3-8B-Base, INFUSER outperforms strong self-evolution baselines with over 20% relative improvement on Olympiad and SuperGPQA benchmarks, and an 8B INFUSER co-evolving generator outperforms a frozen 32B thinking generator on math and coding. Ablations confirm each design choice is necessary, and two extensions, applying INFUSER to an instruction-finetuned anchor and augmenting it with rule-verifiable RLVR data, further demonstrate the flexibility and generalizability of the framework. Code is available at https://github.com/FFishy-git/INFUSER.
Siyu Chen, Miao Lu, Beining Wu +7
Yale University · Stanford University · University of Chicago +2
Self-play supports the self-evolution of language models, but solver performance can plateau or decline across rounds without guidance. Existing unguided methods typically use difficulty, learnability, or diversity signals to keep questions challenging and varied, without identifying which unresolved reasoning weaknesses to target. Existing guided methods rely on external task resources such as human examples, document corpora, or specified difficulty targets. We introduce DiagEvo, which guides question generation using the solver's failure history from self-play, without external task resources. Its diagnostician extracts recurring error causes and stores them in an error-cause memory. The memory groups related causes under skill nodes and tracks each as Active or Mastered according to self-consistency on targeted questions. The challenger uses these states and recurrence counts to balance cause-targeted generation with free exploration. Double-confidence filtering retains intermediate-difficulty questions only when the most common solver answer has a clear vote lead. With the default 4B diagnostician, DiagEvo outperforms all baselines in mean accuracy across nine benchmarks for each solver: Qwen3-4B, Qwen3-8B, and OctoThinker-8B. On Qwen3-8B, DiagEvo reaches 72.3% mean accuracy across five mathematical reasoning benchmarks, 4.5 percentage points above R-Zero. Its overall mean accuracy across nine benchmarks is 57.4%, 3.5 percentage points above SPICE. Ablations show that mixed generation, memory-state updates with cross-state stitching, and double-confidence filtering contribute to these gains.
Xincheng Wei, Yifan Ding, Fucheng Xiong +6
The Chinese University of Hong Kong, Shenzhen · Meituan, LongCat Team · Peking University +2
Self-evolving reasoning frameworks train a Challenger to generate questions exposing a Solver's weaknesses, creating adaptive curricula without human data. However, existing approaches use a single solver's sampling uncertainty as the Challenger's reward. This creates a fundamental bottleneck: as the solver grows confident on the Challenger's question distribution, all sampled answers converge identically, collapsing the reward to zero and starving the Challenger of learning signal. Critically, this single-model reward cannot distinguish genuinely easy questions from those that merely align with one solver's learned biases. We propose a multi-solver disagreement reward using a heterogeneous ensemble varying in model capacity and sampling temperature. A normalized Shannon entropy over the ensemble's per-question plurality answers explicitly rewards questions where solvers produce conflicting solutions---capturing difficulty as inter-model divergence rather than intra-model sampling variance. This richer gradient enables the Challenger to discover questions targeting true capability boundaries, producing a curriculum that forces downstream Solvers to develop robust reasoning strategies generalizing across problem types. Our approach is a drop-in reward function replacement requiring no framework modifications or additional data. Experiments with Qwen3-4B show that Solvers trained on disagreement-Challenger questions achieve +1.34 points average improvement on competition-math benchmarks (MATH-500, AMC, Olympiad), suggesting that multi-solver disagreement provides a complementary and scalable signal for curriculum generation in self-play reasoning systems.
Vinoth Selvendran, Zhanming Zhang
Independent Researcher · Palo Alto, CA, USA · New York, NY, USA