Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient updates and access to expert demonstrations, making it inapplicable at inference. We propose test-time self-distillation, a decoding-time method that extracts a steering signal from the self-distillation framework without any parameter updates, reward models, or training data. Our key insight is that counterfactual contexts, i.e. fixed textual templates that hypothetically prime the model for excellent versus poor reasoning, can substitute for the demonstration. The log-odds ratio of a candidate answer under these two counterfactual conditions defines a new reward signal. We derive the optimal KL-regularized policy under this reward, which takes the form of a Gibbs reweighting of the base distribution. Crucially, this reweighting is global: it cannot be decomposed into independent per-token operations without ignoring future trajectory quality. We therefore approximate the target distribution via beam search. Experiments on mathematical reasoning (MATH500), code generation (HumanEval), and graduate-level science QA (GPQA) across multiple model scales show that test-time self-distillation improves over standard sampling, low temperature, beam search and power sampling baselines on average, demonstrating that the self-distillation principle can be operationalized at inference time.
Figures & tables
Model
Method
MATH500
HumanEval
GPQA
Qwen3.5-0.8B
Base τ=1.0
0.198
0.085
0.020
Low Temperature τ=0.25
0.342
0.225
0.101
Beam Search
0.398
0.207
0.091
Power Sampling τ=0.25
0.430
0.287
0.131
Power Sampling τ=1.0
0.300
0.262
0.060
CBS 1/α=0.25
0.480
0.238
0.106
Table 1: Performance comparison of CBS across 4 models and 3 benchmarks.
Figure 1: Inference efficiency of (a) Qwen2.5-7B and (b) DeepSeek-Math-7B-Instruct on the MATH500 benchmark.
Figure 2: Pass@ k on the MATH500 dataset between Power sampling and CBS (Qwen3.5-2B).
Context
Content
MATH500
No context
-
0.398
Contrastive
Positive: ‘‘This is an example for a response with excellent reasoning:’’ Negative: ‘‘This is an example for a response with wrong reasoning:’’
0.509
Neutral
Positive: ‘‘This is an example for a response to the question assigned to control group one:’’ Negative: ‘‘This is an example for a response to the question assigned to control group two:’’
0.380
Table 2: Context ablation for Qwen3.5-0.8B on MATH500.
Method
Qwen3.5-0.8B
Qwen2.5-7B
Qwen3.5-2B
DeepSeek-Math-7B-Instruct
Beam search ( N=24 )
0.391
0.600
0.676
0.452
CBS 1/α=0.25
0.480
0.610
0.672
0.462
CBS 1/α=0.7
0.436
0.640
0.694
0.458
Table 3: Beam-population budget ablation on MATH500.
Figure 3: Block size K effect on CBS.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Prompt Template
MATH500
System: You are a helpful AI Assistant that provides well-reasoned and detailed responses. You first think about the reasoning process as an internal monologue and then provide the user with the boxed answer. Respond in the following format: <think> … </think> <answer> \boxed{…} </answer>. User: {problem}
HumanEval
System: Please reason step by step internally. Then output ONLY the final Python code that completes the task within triple backticks. Do not include explanations, markdown, or \boxed{}. User: Complete the following Python function: {prompt}
GPQA
System: You are a helpful AI Assistant. Please reason step by step, and put your final answer within \boxed{}. User: Answer the following multiple choice question. The last line of your response should be of the following format: ‘\boxed{$LETTER}’ (without quotes) where LETTER is one of ABCD (ex. ‘\boxed{A}’). Think step by step before answering. {question} A) {choice_A} B) {choice_B} C) {choice_C} D) {choice_D}
Appendix
Table 4: Task Prompt Templates
Role
Serialized Qwen prompt content
System
<|im_start|>system You are a helpful AI Assistant that provides well-reasoned and detailed responses. You first think about the reasoning process as an internal monologue and then provide the user with the boxed answer. Respond in the following format: <think> … </think> <answer> \boxed{…} </answer>. <|im_end|>
User query and positive context
<|im_start|>user Solve for x : 2x+1=32 . This is an example for a response with excellent reasoning: <|im_end|>
Assistant prefix
<|im_start|>assistant
Appendix
Table 5: Example serialization of a MATH500-style query with the positive contrastive context under the Qwen chat template.
Method
Hyperparameters
Standard sampling
N=1 , W=1 , K=32 , τ=1.0
Low-temperature sampling
N=1 , W=1 , K=32 , τ=0.25
Beam search
N=16 , W=4 , K=32 , τ=1.0 , 1/α=0
Contrastive beam search (low α )
N=16 , W=4 , K=32 , τ=1.0 , 1/α=0.25
Contrastive beam search (high α )
N=16 , W=4 , K=32 , τ=1.0 , 1/α=0.7
Power Sampling
Kt=16 , Mt=16 , B=32 (equivalent to K ), τ=0.25 (low) and τ=1.0 (high), α=4
Appendix
Table 6: Main hyperparameter settings for each method.
Figure 4: Pass@ k performance k∈{1,2,3,4} on MATH500 dataset for Power sampling (PS), standard beam search (BS) and contrastive beam search (CBS) using Deepseek-Math-7B-Instruct.
Figure 5: Best-of-16 accuracy for ten contrastive context pairs on MATH500, HumanEval and GPQA (Qwen2.5-7B), at the two deployed steering strengths. Dashed line: no-context baseline.
Content
Contrastive Context Pair
Reasoning
Positive: ‘‘This is an example for a response with excellent reasoning:’’ Negative: ‘‘This is an example for a response with wrong reasoning:’’
Completeness
Positive: ‘‘This is an example for a response with a complete treatment of the problem:’’ Negative: ‘‘This is an example for a response with an incomplete treatment of the problem:’’
Step verification
Positive: ‘‘This is an example for a response that carefully verifies each step of its reasoning:’’ Negative: ‘‘This is an example for a response that skips verification and rushes to a conclusion:’’
Reliability
Positive: ‘‘This is an example for a response that is reliable and trustworthy:’’ Negative: ‘‘This is an example for a response that is unreliable and error-prone:’’
Logical validity
Positive: ‘‘This is an example for a response with logically sound arguments:’’ Negative: ‘‘This is an example for a response with logical fallacies:’’
Reviewer judgment
Positive: ‘‘This is an example for a response that a careful reviewer would approve without hesitation:’’ Negative: ‘‘This is an example for a response that a careful reviewer would reject immediately:’’
Appendix
Table 7: Examples of 10 contrastive context pairs. Every pair shares same initial framing with “This is an example for a response …”. Each positive context is paired with its counterfactual negative.
On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9% and LiveCodeBench v6 pass@12 reaches 70.9%, compared with 69.0% and 66.3% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.
Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecified in long reasoning traces. On-policy self-distillation addresses this by letting the same model act as a teacher conditioned on privileged information, producing a dense per-token signal. But the common choice of a ground-truth answer is only an endpoint cue: on terse-answer tasks, the teacher falls silent at the intermediate positions where path-level guidance matters most. We propose Hindsight Self-Distillation (HSD), which conditions the teacher on a successful peer rollout drawn from the current training group. Such a peer is an exact sample from the success-conditioned policy, requiring no additional sampled rollouts. By providing a full successful continuation rather than only the final answer, the resulting credit signal concentrates at the divergence position between a failed rollout and a successful peer. Across Qwen3-8B and Qwen3-32B on math and code benchmarks, HSD obtains the best result against GRPO variants and on-policy distillation baselines, with the largest gains on terse-answer tasks such as AIME.
Yu Li, Shu Hong, Tian Lan
Department of Electrical and Computer Engineering, George Washington University
On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedback), offers a promising direction for advancing reasoning capability without a stronger external teacher. Yet in math reasoning the gains are inconsistent, even when the same approach succeeds elsewhere. A pointwise mutual information analysis traces the failure to the privileged context itself: it inflates the teacher's confidence on tokens already implied by the solution (structural connectives, verifiable claims) and deflates it on deliberation tokens ("Wait", "Let", "Maybe") that drive multi-step search. We propose Anti-Self-Distillation (AntiSD), which ascends a divergence between student and teacher rather than descending it: this reverses the per-token sign and yields a naturally bounded advantage in one step. An entropy-triggered gate disables the term once the teacher entropy collapses, completing a drop-in replacement for default self-distillation. Across five models from 4B to 30B parameters on math reasoning benchmarks, AntiSD reaches the GRPO baseline's accuracy in 2 to 10x fewer training steps and improves final accuracy by up to 11.5 points. AntiSD opens a path to scalable self-improvement, where a language model bootstraps its own reasoning through its training signal.
Guobin Shen, Xiang Cheng, Chenxiao Zhao +4
1Xiaohongshu Inc. · Institute of Automation, Chinese Academy of Sciences