Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026). Prior work retrieves skills from a bank by semantic relevance, then uses them as inference-time patches or for model distillation. The individual utility of each skill, however, is largely neglected. We first show that, in on-policy distillation where skill-conditioned policies serve as teachers, fewer than 25% of retrieved skills provide useful distillation signals. We then propose SGUID, a method for selecting a compact subset of skills for distillation. SGUID retains a skill only if it consistently yields effective learning signals during training. The selected skills are then distilled to produce a better model. Our results show that not all skills are worth distilling. Across four models from the Olmo and Qwen families, distilling 6 selected skills matches or exceeds full-bank distillation in mean avg@12 on three of the four models, and on all four after a second round that distills 3 newly selected skills, while the full banks are up to 11x larger. Importantly, SGUID supports stable model-skill co-evolution: after a distillation round, a new candidate bank is curated from the updated model's rollouts, and SGUID selects which skills to internalize next. In the second round, this loop selects 3 new skills and improves Qwen3-8B from 64.3% to 66.3%. The selection step is essential for stability: on Qwen3-4B, naively updating the model with unfiltered skills degrades performance, including a 0.3 percentage point drop on HMMT25, whereas SGUID improves HMMT25 by 0.5 points after the first round and 1.1 points after the second. These results identify skill selection as the key mechanism for stable model-skill co-evolution.
Figures & tables
Figure 1: Overview of SGUID at co-evolution round r . (1) Curate: the current policy θr rolls out on training problems, the rollouts are checked against the reference solutions y⋆ to curate a candidate skill bank Bθr (the bank size shown is that of Qwen3-8B ). (2) Select skills by training signal: the student PS(⋅∣x) , initialized from θr , generates a rollout y^ that is scored by privileged teachers PT(⋅∣x,si) , each conditioned on a retrieved skill si . Skills that contribute early, still produce at least h positive events late, and whose contribution rate decreases by at most a factor of τ are kept; skills whose signal decays too fast, stays at zero, or that are never retrieved are discarded. (3) Internalize: restarting from θr , the compact bank Bθr with ∣Bθr∣=K≪∣Bθr∣ is distilled into the model, producing θr+1 , which becomes the policy for the next round.
Figure 2: Skill retrieval does not guarantee active teacher signals. <25% of skills are active.
Model
Method
# Distilled Skills
AIME24
AIME25
HMMT25
Mean
Olmo-3-7B
Base
–
53.1
39.4
26.9
39.8
GRPO
–
53.6
42.5
25.8
40.6
OPSD
–
42.5
37.5
20.8
33.6
SGSD
38
54.4
41.1
24.4
40.0
SGUID , R1
6
55.0
42.2
25.6
40.9
SGUID , R2
3
55.0
42.8
26.1
41.3
Table 1: Main results across three math reasoning benchmarks. We report the avg@12 accuracy for the best checkpoint according to the average performance on AIME24, AIME25, and HMMT25 across 200 training steps. Bold marks the best and underline the second best. # Distilled Skills is the number of skills the teacher can retrieve from; for SGSD we follow the original banks in Huang et al. (2026a) . R1 and R2 denote the first and second co-evolution rounds of SGUID (ours). The skills selected for each model and round are listed in Appendix B .
Figure 3: Two rounds of model–skill co-evolution on Qwen3-4B . Without selection (distilling the full bank), the mean gain shrinks from round one to round two; in contrast, SGUID (w/ selection) improves the model steadily over rounds.
Model
Method
AIME24
AIME25
HMMT25
Mean
Qwen3-1.7B
Base
50.6
35.6
23.1
36.4
full, 200 steps
55.3
43.1
28.6
42.3
25 steps
56.1
43.9
27.5
42.5
Qwen3-4B
Base
73.9
65.8
46.7
62.1
full, 200 steps
76.7
71.4
47.2
65.1
25 steps
72.5
70.8
46.9
63.4
Table 2: Qwen3-1.7B and Qwen3-4B avg@12 when the compact bank is selected using training statistics from the full 200-step run ( SGUID (full)) or only the first 25 steps ( SGUID (25 steps)). The 25-step proxy matches the full run on 1.7B but is less reliable on 4B.
Method
AIME24
AIME25
HMMT25
Mean
Base
73.9
65.8
46.7
62.1
OPSD
77.8
68.9
46.1
64.3
SGSD
75.6
70.3
46.7
64.2
K=2
76.9
69.4
43.6
63.3
K=4
75.6
70.8
46.9
64.4
K=6
76.7
71.4
47.2
65.1
Table 3: How compact should the bank be? Qwen3-4B avg@12 as the number of retained skills K varies. We use K=6 by default in the first round. The last row keeps the same budget but ranks skills by AL(s) alone, dropping the persistence constraint QL(s)≥QE(s)/τ .
Method
AIME24
AIME25
HMMT25
Mean
Base
50.6
35.6
23.1
36.4
SGSD
55.6
46.1
27.2
43.0
top-6
55.3
43.1
28.6
42.3
random-6
54.1 ± 0.8
42.6 ± 1.2
26.9 ± 0.9
41.2 ± 0.9
worst-6
55.6
42.8
26.1
41.5
Table 4: Qwen3-1.7B avg@12 with K=6 held fixed, varying only which 6 skills are retained: the top-ranked, random (over 3 seeds for skill selection), or bottom-ranked (worst) 6 according to our selection signal. All skill selection baselines are trained for 100 steps. Bold marks the best value among the three K=6 variants; Base and SGSD are reference rows.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
GRPO
OPSD
SGSD
SGUID
Learning rate
5×10−6
5×10−6
5×10−6
5×10−6
Completion length
16000
1024
1024
1024
Prompt length
2048
20000
20000
20000
Temperature
1.2
1.1
1.1
1.1
Top- p / top- k
–
0.95 / 20
0.95 / 20
0.95 / 20
Generations per prompt
8
1
1
1
Appendix
Table 5: Training hyperparameters for each method. LoRA is applied to q_proj , k_proj , v_proj , o_proj , gate_proj , up_proj and down_proj in all runs. Dashes mark settings that do not apply.
ID
Title
Principle
Round 1 ( K=6 )
gen_007
Solve triangle inequalities
Ensure sequences maintain triangle inequality conditions by enforcing each term < sum of two preceding terms.
gen_017
Balanced parity counting
When the number of even and odd elements are equal, the minimum number of swaps required is half the total number of elements.
gen_019
Count with constraints
Calculate the number of valid configurations under specified constraints using combinatorial methods or factor pair analysis.
gen_021
Integer factor pair analysis
Identify valid integer factor pairs for area or perimeter constraints to determine dimensions.
gen_023
Count under constraints with digit analysis
Calculate the number of integers within a range that meet specific conditions, including multiples and digit patterns.
Appendix
Table 6: Compact skill banks selected for Qwen3-1.7B .
ID
Title
Principle
Round 1 ( K=6 )
gen_002
Simplify Expressions
Decompose complex expressions or ratios into simpler components using algebraic techniques, factoring, or number theory to reduce complexity.
gen_001
Translate Constraints to Algebra
Convert geometric, number-theoretic, or combinatorial conditions into algebraic equations or expressions to isolate variables or derive relationships.
gen_006
Analyze Symmetries
Adjust counts or configurations by dividing total possibilities by symmetry group order to account for equivalent arrangements.
gen_010
Constraint Optimization
Maximize/minimize expressions or narrow down cases by analyzing variable relationships, parity, or integer properties under constraints.
gen_007
Validate Solutions
Test candidate solutions against original constraints or edge cases to ensure consistency across scenarios.
Appendix
Table 7: Compact skill banks selected for Qwen3-4B .
ID
Title
Principle
Round 1 ( K=6 )
gen_045
Geometric Constraint Modeling
Formulate algebraic equations from geometric relationships and properties.
gen_009
Number Theory
Apply divisibility rules, modular arithmetic, and prime factorization to analyze number properties and solve equations.
gen_046
Tangency Relationship Derivation
Use geometric perpendicularity properties to derive solution relationships.
gen_024
Combinatorial Counting
Systematically count configurations by iterating through constraints or categorizing arrangements.
gen_039
Integer Scaling Optimization
Minimize target quantities by testing minimal integer scaling parameters with common factors.
Appendix
Table 8: Compact skill banks selected for Qwen3-8B .
ID
Title
Principle
Round 1 ( K=6 )
gen_027
Diophantine Analysis with Parameterization
Analyze integer solutions by fixing parameters and applying Diophantine conditions.
gen_010
Translate Verbal or Qualitative Problems to Mathematics
Convert non-technical or narrative problem statements into mathematical models or equations.
gen_002
Simplify Expressions Using Algebra and Number Theory
Apply algebraic and number-theoretic techniques to reduce expression complexity.
gen_007
Evaluate and Sum Series
Apply closed-form formulas to compute or analyze arithmetic, geometric, or power series.
gen_022
Geometric Mean and Right Triangle Properties
Apply geometric mean properties of inscribed squares in right triangles to find areas or lengths.
Appendix
Table 9: Compact skill banks selected for Olmo-3-7B .
State Key Lab of CAD&CG, Zhejiang University · School of Software and Microelectronics, Peking University · School of Software Technology, Zhejiang University +1