Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.
Figures & tables
Figure 1: A distillation attack’s pipeline. Attackers amass large volumes of reasoning traces from proprietary LLMs and train (“distill”) their own models on these traces, followed by further training of distilled models with reinforcement learning (RL). Existing threat models assume attackers train models only with distillation, while a realistic distillation attack likely follows distillation with RL.
Figure 2: RL outperforms distillation, with distillation followed by RL outperforming both. An attacker aiming to train a state-of-the-art model would likely use distillation to bootstrap subsequent RL, rather than relying solely on distillation. Accuracies are averages over the GSM8K ( Cobbe et al., 2021 ) , Minerva Math ( Hendrycks et al., 2021 ) , and MATH500 ( Lightman et al., 2024 ) datasets.
Figure 3: Distillation improves pass@ k for high k , allowing futher RL to lead to accuracy improvements (pass@1) . In this sense, distillation bootstraps subsequent RL training. Pass@ k is averaged over the GSM8K, Minerva, and MATH500 datasets. A detailed breakdown of results is available in Appendix A.2 .
Figure 4: Relative to a perfect teacher , RL (dashed) is mode-seeking while distillation (dotted) is mode-covering. Green regions represent correct answers, yellow represents incorrect. Probability densities are presented unnormalized. In practice, RL is more limited by the initial distribution of the base model being trained than distillation.
Figure 5: Antidistillation sampling can be effective following distillation but break following RL. Antidistillation sampling modifies teacher traces to “poison” distilled students, degrading their performance after distillation. A Qwen2.5-0.5B student is distilled normally and with low, mild, or high poisoning levels; after RL, low- to mildly poisoned student models close the performance gap with the unpoisoned model. An attacker using RL would benefit similarly from antidistillation-sampled data as from regular data. Left: the average accuracy over the GSM8K, MATH500, and Minerva datasets. Right: the performance over Minerva, which is harder than GSM8K and similar in difficulty to MATH500. Additional details are provided in Appendix B .
Figure 6: A simple attack to approximately reconstruct closed-source models’ reasoning traces. Note that the reconstructed traces need not be faithful to the hidden, typically unknown, full reasoning traces, but only similarly useful for bootstrapping a model’s reasoning.
Figure 7: Summaries leak sufficient information to distill reasoning capabilities equal to those achieved using full traces. After reinforcement learning, the model distilled on expanded summaries (right) performs similarly to the model distilled on full, unobfuscated reasoning traces (middle). Distilling on either full traces or expanded summaries decently outperforms no distillation (left). Open-source setting, with Qwen2.5-14B-RL as the attacked teacher, and Llama-3.2-3B base as the attacker’s model.
Figure 8: Summaries leak sufficient information to distill reasoning capabilities from closed-source models. Full traces were obtained with the extraction attack of Panfilov et al. (2026) (Appendix C.2 ), excluding Gemini models, which had the extraction attack patched at the time of writing. After reinforcement learning, a base model distilled on expanded summaries performs similarly to the same model distilled on full traces. Expanded summaries are constructed using information readily available through the model APIs.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Model
GSM8K
Minerva
Math500
Average
Qwen2.5-0.5B
Base
35.6%
18.2%
19.2%
24.3%
Base-RL
48.5%
36.9%
35.4%
40.3%
Distill
45.1%
30.7%
31.6%
35.8%
Distill-RL
49.1%
40.2%
35.4%
41.6%
Appendix
Table 1: Per-dataset breakdown for results in Figure 2 . RL outperforms distillation, with distillation followed by RL outperforming only distillation and only RL.
Figure 9: Per-dataset pass@ k curves are qualitaitvely similar to their average, shown in Figure 3 . Results are shown for Llama-3.2-3B (left), Qwen2.5-3B (middle), and Gemma-2B (right), across the GSM8K (top), Minerva (middle), and MATH500 (bottom) datasets.
Figure 10: The Qwen2.5-Math-7B primarily reasons in Python code and hallucinates its corresponding outputs. This is illustrated with sample reasoning traces of both a correct (left) and an incorrect (right) answer. Traces are left in their raw form for illustration, with some code and reasoning being omitted.
Figure 11: RL training fails to eliminate the undesirable pattern of reasoning with Python code when it is the dominant approach used by a base model. The Qwen2.5-7B base model (top) rarely answers questions correctly using code, so RL training eliminates its code-reasoning behavior. In contrast, Qwen2.5-Math-7B (middle) obtains correct answers predominantly with code, and RL training through two frameworks ( Liu et al., 2025a ; Zeng et al., 2025 ) reinforces this code-reasoning behavior.
Dataset
Model
GSM8K
Minerva
Math500
Average
Qwen2.5-0.5B
Base
35.6%
18.2%
19.2%
24.3%
Base-RL
48.5%
36.9%
35.4%
40.3 %
Distill
47.0%
31.5%
31.4%
36.6%
Qwen2.5-1.5B
Appendix
Table 2: Distillation with rejection sampling still underperforms reinforcement learning. Rejection sampling is done by allowing the teacher 8 attempts at every question, and distillation is then done only over correct-answer traces.
Dataset
Change
GSM8K
Minerva
Math500
Average
None
89.0%
75.9%
75.4%
80.1 %
Qwen2.5-7B-RL Teacher
88.9%
75.9%
76.4%
80.4%
DeepSeek-R1-Distill-Qwen-14B Teacher
87.4%
71.5%
69.4%
76.1%
Qwen3-14B Teacher
83.5%
41.5%
38.2%
54.4%
s1K Dataset
88.7%
74.2%
75.8%
79.6%
Appendix
Table 3: Further distillation ablations on Qwen2.5-7B do not qualitatively change results. Changing the default distillation setup (see Section A.1 ) to use a smaller teacher (Qwen2.5-7B-RL ( Zeng et al., 2025 ) ) marginally improves performance. Changing the setup either to use a different teacher (Deepseek-R1-Distill-Qwen-14B ( Guo et al., 2025a ) or Qwen3-14B ( Yang et al., 2025 ) ), or a different training dataset (s1K ( Muennighoff et al., 2025 ) ) lowers distillation performance.
Description
GSM8K
Reference
Base
35.6%
RL
48.5 %
Distilled
45.1%
Training Dataset
SimpleRL-Zoo (Hard)
47.4%
Appendix
Table 4: Further distillation ablations on Qwen2.5-0.5B do not qualitatively change results . Performance improves only when using the harder SimpleRLZoo dataset ( Zeng et al., 2025 ) to generate the teacher’s traces, with no gains from using the s1K ( Muennighoff et al., 2025 ) dataset, or 1000 questions from AIME ( Veeraboina, 2023 ) . All other ablations also use the hard SimpleRLZoo dataset, but do not yield significant performance improvements.
Dataset
Model
GSM8K
Minerva
Math500
Average
DAPO-Qwen2.5-32B (RL)
94.4 %
94.7 %
91.6 %
93.6 %
Qwen2.5-32B-Distill-DeepseekR1
93.1%
91.0%
90.2%
91.4%
Qwen2.5-32B-SimpleRLZoo
85.5%
85.4%
84.4%
85.1%
Appendix
Table 5: Even at larger scales, RL outperforms distillation at improving a base model’s reasoning accuracy. Qwen2.5-32B base models trained with RL and distillation using a harder dataset. The 32B SimpleRLZoo model underperforms, likely due to using too easy questions in RL training. Open-source versions of each model were evaluated using the same evaluation framework as models at the 7B scale (see Appendix A.1 .
Dataset
Model
GSM8K
Minerva
Math500
Average
Before RL
Base
35.6%
18.2%
19.2%
24.3%
No Poison
44.7%
33.4%
36.0%
38.0%
Low Poisoning
43.5%
31.0%
30.2%
34.9%
Mild Poisoning
40.9%
25.9%
25.4%
30.7%
Appendix
Table 6: Per-dataset for results in Figure 5 . After RL, Qwen2.5-0.5B base models distilled with low or mild poisoning match the unpoisoned distilled model’s performance. High poisoning lowers post-RL performance.
Dataset
Model
GSM8K
Minerva
Math500
Average
Before RL
Base
21.1%
7.3%
7.8%
12.1%
λ=0
30.5%
12.8%
12.2%
18.5%
λ=0.01
30.6%
12.5%
12.2%
18.4%
λ=0.02
32.2%
12.9%
11.0%
18.7%
Appendix
Table 7: Evaluation breakdown for antidistillation on Llama-3.2-3B. λ≥0.03 is required to create performance degradations after distillation, requiring a minimum loss of 10% in relative teacher accuracy. However, results indicate that antidistillation sampling can in some cases be effective.
Figure 12: Extracted reasoning traces match the length reported by the provider. Each point denotes a single SimpleRL-Zoo trainining set question. Extracted and reported lengths differ by at most 1% for 99.9% of traces for GPT-5 mini (within one 64-token bin, the API’s rounding), 99.9% for Claude Sonnet 4.6 with a thinking budget of 8192 tokens, and 99.4% for Claude Sonnet 4.6 at maximum reasoning effort.
Figure 13: Performance improvements from expanded reasoning traces come from semantic information in summaries, not the expander model. After reinforcement learning, distilling on the traces generated from the expander model leads to no performance gains (middle left), with performance equivalent to RL with no distillation (left). Results with full traces (middle right) and expanded summaries (right) are included for reference.
Figure 21
Dataset
Llama-3.2-3B
GSM8K
Minerva
Math500
Average
Base
0.8%
3.1%
2.6%
2.2%
Base-RL
40.7%
17.3%
14.8%
24.3%
Round 1: Distillation on Medium Difficulty Data
Full Traces
33.4%
13.7%
12.4%
19.8%
Expanded Traces
24.0%
10.3%
10.4%
14.9%
Appendix
Table 8: A second round of training with distillation and reinforcement further improves performance. Further, after a second round of distillation and RL, the model trained on expanded traces recovers lost performance on GSM8K. Training uses the medium and hard datasets from Zeng et al. (2025) .
Figure 14: Closed-source traces obtained with higher effort lead to marginally better performance . In evaluations after further reinforcement learning, distilling on full traces extracted from Claude Sonnet 4.6 (right) yields higher performance than using a thinking budget of 8192 tokens (left). Distilling on expanded summaries (middle) is included for reference.
Dataset
Model
GSM8K
Minerva
Math500
Average
Distillation with Forced Thinking
GPT-5 mini
35.0%
12.3%
12.4%
19.9%
Claude Sonnet 4.6
33.5%
12.8%
12.0%
19.4%
+ Reinforcement Learning
GPT-5 mini
47.2%
21.9%
17.2%
28.8%
Appendix
Table 9: Distillation on full traces from closed-source models with forced thinking improves performance. Distillation is done with the Llama-3.2-3B base model on full traces from closed-source models with (left) and without (right) forced thinking, as described in Appendix C.1 .
Figure 15: Trace expansion is necessary. After reinforcement learning, distilling the Llama-3.2-3B base model on open-source summaries alone leads to no performance gains (middle left): performance is essentially equivalent to RL with no distillation (left). Results with full traces (middle right) and expanded summaries (right) are included for reference.
Figure 16: Qualitative examples of raw Qwen2.5-7B-RL traces sampled using antidistillation sampling with increasing poisoning levels. Poisoned traces are shown given no (top), mild (middle) and high (bottom) poisoning. All traces are answers to the following level 3 math question from the hard SimpleRLZoo dataset ( Zeng et al., 2025 ) : Uri buys two burgers and a soda for \2.10,andGenbuysaburgerandtwosodasfor$2.40$ . How many cents does a soda cost?
Figure 17: A final answer and summary in the open-source setting. Outputs are shown for the same level 3 math question from the medium SimpleRLZoo dataset ( Zeng et al., 2025 ) . Top: the question and the answer from the teacher model, Qwen2.5-14B-RL. Bottom: the summary of the final answer generated with Qwen2.5-7B-Instruct Qwen et al. (2025) . Traces are formatted for readability.
Figure 18: Examples of final answers and reasoning summaries exposed by current APIs. Outputs are shown for the same level 3 math question from the medium SimpleRLZoo dataset ( Zeng et al., 2025 ) . Top: the question and the final answer from GPT-5 mini. Middle: GPT-5 mini’s reasoning summary. Bottom: Claude Sonnet 4.6’s summarized thinking for the same question. Traces are formatted for readability.
Figure 19: System prompt used to summarize full reasoning traces from an open-source model.
Figure 20: Prompt used to expand summarized reasoning traces.
Figure 21: Prompt used to synthesize summaries and final answers into a detailed breakdown of reasoning steps.