A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.
Figures & tables
Figure 1: Overview of AdviSD . (A) Multi-turn advising. Before each executor response, the advisor provides advice or abstains; the frozen executor uses its native context and any issued advice. (B) Training from targeted feedback. Reflection proposes corrections at advice decisions in imperfect episodes. Among flagged decisions, AdviSD retains original abstentions and those whose scores with and without issued advice differ by more than a calibrated threshold. The pre-update advisor computes both on the same recorded executor response. At retained decisions, a feedback-conditioned copy of the pre-update advisor teaches the trainable advisor, which sees only the original context. This self-distillation supplements GRPO on all episodes.
EnvScaler
BFCL-v3
Adaptation method
Score
Base
Miss Func
Miss Param
Long Ctx
Avg
Frozen executor: Gemini 3.7 Flash
Standalone executor: no advisor
No advisor
75.7±2.0
70.8±1.9
60.4±1.4
42.9±1.9
57.1±1.9
57.8±1.4
Executor prompt optimization: no advisor
GEPA ( Agrawal et al., 2026a )
76.0±1.8
69.6±1.9
61.7±1.4
45.4±1.4
60.8±1.4
59.4±0.3
Table 1: In-domain test performance. EnvScaler: native score ×100 ; BFCL-v3: official-checker accuracy (%) on 320 held-out tasks; Avg weights categories equally. Mean ± sample SD follows Section 7 . Bold/underline: largest/second-largest distinct means, including ties, per column/executor.
BFCL-v3
EnvScaler
Advisor training
Gemini 3.7 Flash
Claude Sonnet 4.6
Gemini 3.7 Flash
Claude Sonnet 4.6
GRPO
61.4±1.0
58.9±0.7
79.7±2.2
80.7±2.3
AdviSD , no gate
62.7±0.8
59.7±0.3
81.1±1.5
81.4±0.9
AdviSD , inverted gate
61.1±1.4
58.3±0.6
78.7±1.9
79.4±1.6
AdviSD , matched-count random
62.9±0.9
60.6±0.4
80.9±1.3
81.1±1.2
AdviSD
67.8±1.6
63.1±0.6
84.8±1.6
84.6±1.1
Table 2: Correction selection and abstention supervision (Q2). All AdviSD variants share GRPO, reflection, teacher construction, and the auxiliary loss and weight schedule. Self-distillation uses feasible reflection-flagged decisions. AdviSD retains issued-advice decisions whose with- and without-advice scores differ by more than the threshold in either direction; inverted gate retains those with absolute score differences at or below the threshold, without count matching. No gate retains every proposal. Matched-count random samples issued-advice decisions uniformly, matching the gate’s per-episode count on the random control’s own rollouts. All four retain flagged original abstentions. No abstention bypass removes self-distillation at original abstentions while keeping the ordinary gate, GRPO, and the advisor’s ability to abstain. GRPO alone uses no self-distillation. Metrics and averaging follow Table 1 ; bold marks column maxima.
ACEBench
ToolHop
τ2 -bench
RoTBench
Adaptation method
M-Step
M-Turn
AC
Airline
Retail
Telecom
TS
PI
CF
Macro Avg
Frozen executor: Gemini 3.7 Flash
Standalone executor: no advisor
No advisor
86.7±2.9
66.7±5.8
81.3±0.4
84.0±2.0
58.2±1.3
91.5±1.3
71.9±0.8
66.1±0.4
47.8±0.2
74.5
Executor prompt optimization: no advisor
GEPA
85.0±2.2
63.9±2.1
80.6±0.1
80.0±3.5
55.8±1.3
88.6±0.9
68.8±0.4
63.4±0.3
45.9±0.2
72.3
Table 3: Out-of-domain (OOD) performance without retraining (Q3). BFCL-selected systems; mean ± sample SD follows the evaluation protocol . Metrics (%): ACEBench success, ToolHop answer correctness (AC), τ2 -bench pass 1 , and RoTBench tool selection (TS), parameter identification (PI), and content filling (CF). Macro Avg averages components within each benchmark, then the four benchmarks equally (Appendix E ). Bold/underline mark the largest/second-largest distinct means per column and executor, including ties.
2 task groups per minibatch; microbatch 1 per GPU; 1 update epoch
Learning rate / reference penalty
10−6 / 0.001 ; fixed reference policy
Appendix
Table 4: BFCL training and evaluation settings. Token limits are configured capacities; a rollout need not reach them.
Benchmark
Split or evaluation inventory
Reported metric / protocol
BFCL
400 train, 80 validation, 320 test; four equally sized categories
Official backend-state and execution-response success
EnvScaler
1,880 train, 470 validation, 200 test
100× mean native fractional task score
ACEBench
20 multi-step and 30 multi-turn identities
End-to-end success; native tools and user simulation
ToolHop
995 queries; 3,912 tools
Answer Correctness; Free protocol, at most 9 executor responses
τ2 -bench
50 airline, 114 retail, 114 telecom entries
Per-domain pass 1 ; native user simulator and policies
RoTBench
840 records from 105 questions: 105 Clean, 210 each Slight/Medium/Heavy, 105 Union
Released TS, PI, CF; first-turn text prediction, without executing tools
Appendix
Table 5: Datasets and primary metrics. External counts describe the reference evaluation inventories; repeats do not create new task identities.
Metric
Early
Middle
Late
Issued abstention (%)
23.8
39.3
43.5
Proposals per reflected episode
4.2
2.2
2.0
Ordinary gate retention (%)
58.0
36.9
32.2
Abstention-bypass share (%)
13.9
19.9
20.3
Supervised decisions per rollout episode
1.7
0.5
0.4
Episode coverage (%)
60.2
35.5
31.3
Appendix
Table 6: Evolution of targeted supervision on BFCL-v3. Statistics describe one training run with frozen Claude Sonnet 4.6 and run-specific threshold ϵc=0.411632 . Early, middle, and late refer to updates 1–20, 21–100, and 101–200. Definitions below distinguish advisor decisions, reflection proposals, and rollout episodes.
Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this improvement when the context is not present. The method works by matching the model's output distribution under two settings: a student that sees only the question, and a self-teacher that also sees the context. What the model learns therefore depends on what context the self-teacher receives, yet the design of this context remains largely unexplored. We study context design for self-distillation by training a solver on feedback from a frozen critic. We compare three conditions: (i) a binary reward (GRPO), (ii) the reference solution, and (iii) a step-by-step critique aligned to the solver's reasoning trace. Step-aligned critique yields the largest gains, outperforming GRPO by 16.11 points and reference-solution-conditioned self-distillation by 5.27 points (Avg@12). Per-token advantage analysis reveals why: step-aligned feedback targets only the tokens where reasoning fails, leaving correct behavior intact. Conditioning on the reference solution, by contrast, pressures the model to change its behavior at every token (even correct steps) because an alternative derivation inevitably differs in phrasing and approach. This suggests that structural alignment between feedback and the solver's reasoning is a key driver of self-distillation effectiveness.
Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teacher but typically assign it a fixed role: direct distribution matching may destabilize successful behavior, while magnitude-only modulation offers little corrective guidance after failure. We observe that successful and failed trajectories require different forms of hindsight supervision. A successful response already contains a valid student-generated reasoning path and can therefore serve as privileged context rather than being replaced by an external rationale. A failed response, however, requires corrective reference information. We introduce Hybrid Hindsight Self-Distillation (H2SD), which jointly adapts teacher context and update strategy to trajectory correctness. For successful trajectories, we construct the teacher context from the verified response and a rephrasing instruction, and use the teacher only to re-evaluate the original response tokens. The resulting probabilities refine token credit assignment without changing the direction determined by the reward. For failed trajectories, a verified reference hint provides corrective guidance through reverse-KL distillation. Experiments on challenging reasoning benchmarks show that H2SD achieves the strongest overall performance among representative RLVR and self-distillation baselines, with stable optimization and a favorable accuracy-efficiency trade-off.
Qiye Cai, Yichuan Ma, Peiji Li +6
1Shanghai Artificial Intelligence Laboratory · 2Harbin Institute of Technology · 3Fudan University +1
Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at https://github.com/zhansan114514/OPSRD.
Weijie Ren, Yanwen Zhang, Hao Li +3
Zhejiang University · University of Electronic Science and Technology of China · University of Science and Technology of China