A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.
Figures & tables
Figure 1: Overview of AdviSD . (A) Multi-turn advising. Before each executor response, the advisor provides advice or abstains; the frozen executor uses its native context and any issued advice. (B) Training from targeted feedback. Reflection proposes corrections at advice decisions in imperfect episodes. Among flagged decisions, AdviSD retains original abstentions and those whose scores with and without issued advice differ by more than a calibrated threshold. The pre-update advisor computes both on the same recorded executor response. At retained decisions, a feedback-conditioned copy of the pre-update advisor teaches the trainable advisor, which sees only the original context. This self-distillation supplements GRPO on all episodes.
EnvScaler
BFCL-v3
Adaptation method
Score
Base
Miss Func
Miss Param
Long Ctx
Avg
Frozen executor: Gemini 3.7 Flash
Standalone executor: no advisor
No advisor
75.7±2.0
70.8±1.9
60.4±1.4
42.9±1.9
57.1±1.9
57.8±1.4
Executor prompt optimization: no advisor
GEPA ( Agrawal et al., 2026a )
76.0±1.8
69.6±1.9
61.7±1.4
45.4±1.4
60.8±1.4
59.4±0.3
Table 1: In-domain test performance. EnvScaler: native score ×100 ; BFCL-v3: official-checker accuracy (%) on 320 held-out tasks; Avg weights categories equally. Mean ± sample SD follows Section 7 . Bold/underline: largest/second-largest distinct means, including ties, per column/executor.
BFCL-v3
EnvScaler
Advisor training
Gemini 3.7 Flash
Claude Sonnet 4.6
Gemini 3.7 Flash
Claude Sonnet 4.6
GRPO
61.4±1.0
58.9±0.7
79.7±2.2
80.7±2.3
AdviSD , no gate
62.7±0.8
59.7±0.3
81.1±1.5
81.4±0.9
AdviSD , inverted gate
61.1±1.4
58.3±0.6
78.7±1.9
79.4±1.6
AdviSD , matched-count random
62.9±0.9
60.6±0.4
80.9±1.3
81.1±1.2
AdviSD
67.8±1.6
63.1±0.6
84.8±1.6
84.6±1.1
Table 2: Correction selection and abstention supervision (Q2). All AdviSD variants share GRPO, reflection, teacher construction, and the auxiliary loss and weight schedule. Self-distillation uses feasible reflection-flagged decisions. AdviSD retains issued-advice decisions whose with- and without-advice scores differ by more than the threshold in either direction; inverted gate retains those with absolute score differences at or below the threshold, without count matching. No gate retains every proposal. Matched-count random samples issued-advice decisions uniformly, matching the gate’s per-episode count on the random control’s own rollouts. All four retain flagged original abstentions. No abstention bypass removes self-distillation at original abstentions while keeping the ordinary gate, GRPO, and the advisor’s ability to abstain. GRPO alone uses no self-distillation. Metrics and averaging follow Table 1 ; bold marks column maxima.
ACEBench
ToolHop
τ2 -bench
RoTBench
Adaptation method
M-Step
M-Turn
AC
Airline
Retail
Telecom
TS
PI
CF
Macro Avg
Frozen executor: Gemini 3.7 Flash
Standalone executor: no advisor
No advisor
86.7±2.9
66.7±5.8
81.3±0.4
84.0±2.0
58.2±1.3
91.5±1.3
71.9±0.8
66.1±0.4
47.8±0.2
74.5
Executor prompt optimization: no advisor
GEPA
85.0±2.2
63.9±2.1
80.6±0.1
80.0±3.5
55.8±1.3
88.6±0.9
68.8±0.4
63.4±0.3
45.9±0.2
72.3
Table 3: Out-of-domain (OOD) performance without retraining (Q3). BFCL-selected systems; mean ± sample SD follows the evaluation protocol . Metrics (%): ACEBench success, ToolHop answer correctness (AC), τ2 -bench pass 1 , and RoTBench tool selection (TS), parameter identification (PI), and content filling (CF). Macro Avg averages components within each benchmark, then the four benchmarks equally (Appendix E ). Bold/underline mark the largest/second-largest distinct means per column and executor, including ties.
2 task groups per minibatch; microbatch 1 per GPU; 1 update epoch
Learning rate / reference penalty
10−6 / 0.001 ; fixed reference policy
Appendix
Table 4: BFCL training and evaluation settings. Token limits are configured capacities; a rollout need not reach them.
Benchmark
Split or evaluation inventory
Reported metric / protocol
BFCL
400 train, 80 validation, 320 test; four equally sized categories
Official backend-state and execution-response success
EnvScaler
1,880 train, 470 validation, 200 test
100× mean native fractional task score
ACEBench
20 multi-step and 30 multi-turn identities
End-to-end success; native tools and user simulation
ToolHop
995 queries; 3,912 tools
Answer Correctness; Free protocol, at most 9 executor responses
τ2 -bench
50 airline, 114 retail, 114 telecom entries
Per-domain pass 1 ; native user simulator and policies
RoTBench
840 records from 105 questions: 105 Clean, 210 each Slight/Medium/Heavy, 105 Union
Released TS, PI, CF; first-turn text prediction, without executing tools
Appendix
Table 5: Datasets and primary metrics. External counts describe the reference evaluation inventories; repeats do not create new task identities.
Metric
Early
Middle
Late
Issued abstention (%)
23.8
39.3
43.5
Proposals per reflected episode
4.2
2.2
2.0
Ordinary gate retention (%)
58.0
36.9
32.2
Abstention-bypass share (%)
13.9
19.9
20.3
Supervised decisions per rollout episode
1.7
0.5
0.4
Episode coverage (%)
60.2
35.5
31.3
Appendix
Table 6: Evolution of targeted supervision on BFCL-v3. Statistics describe one training run with frozen Claude Sonnet 4.6 and run-specific threshold ϵc=0.411632 . Early, middle, and late refer to updates 1–20, 21–100, and 101–200. Definitions below distinguish advisor decisions, reflection proposals, and rollout episodes.