Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient updates and access to expert demonstrations, making it inapplicable at inference. We propose test-time self-distillation, a decoding-time method that extracts a steering signal from the self-distillation framework without any parameter updates, reward models, or training data. Our key insight is that counterfactual contexts, i.e. fixed textual templates that hypothetically prime the model for excellent versus poor reasoning, can substitute for the demonstration. The log-odds ratio of a candidate answer under these two counterfactual conditions defines a new reward signal. We derive the optimal KL-regularized policy under this reward, which takes the form of a Gibbs reweighting of the base distribution. Crucially, this reweighting is global: it cannot be decomposed into independent per-token operations without ignoring future trajectory quality. We therefore approximate the target distribution via beam search. Experiments on mathematical reasoning (MATH500), code generation (HumanEval), and graduate-level science QA (GPQA) across multiple model scales show that test-time self-distillation improves over standard sampling, low temperature, beam search and power sampling baselines on average, demonstrating that the self-distillation principle can be operationalized at inference time.
Figures & tables
Model
Method
MATH500
HumanEval
GPQA
Qwen3.5-0.8B
Base τ=1.0
0.198
0.085
0.020
Low Temperature τ=0.25
0.342
0.225
0.101
Beam Search
0.398
0.207
0.091
Power Sampling τ=0.25
0.430
0.287
0.131
Power Sampling τ=1.0
0.300
0.262
0.060
CBS 1/α=0.25
0.480
0.238
0.106
Table 1: Performance comparison of CBS across 4 models and 3 benchmarks.
Figure 1: Inference efficiency of (a) Qwen2.5-7B and (b) DeepSeek-Math-7B-Instruct on the MATH500 benchmark.
Figure 2: Pass@ k on the MATH500 dataset between Power sampling and CBS (Qwen3.5-2B).
Context
Content
MATH500
No context
-
0.398
Contrastive
Positive: ‘‘This is an example for a response with excellent reasoning:’’ Negative: ‘‘This is an example for a response with wrong reasoning:’’
0.509
Neutral
Positive: ‘‘This is an example for a response to the question assigned to control group one:’’ Negative: ‘‘This is an example for a response to the question assigned to control group two:’’
0.380
Table 2: Context ablation for Qwen3.5-0.8B on MATH500.
Method
Qwen3.5-0.8B
Qwen2.5-7B
Qwen3.5-2B
DeepSeek-Math-7B-Instruct
Beam search ( N=24 )
0.391
0.600
0.676
0.452
CBS 1/α=0.25
0.480
0.610
0.672
0.462
CBS 1/α=0.7
0.436
0.640
0.694
0.458
Table 3: Beam-population budget ablation on MATH500.
Figure 3: Block size K effect on CBS.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Prompt Template
MATH500
System: You are a helpful AI Assistant that provides well-reasoned and detailed responses. You first think about the reasoning process as an internal monologue and then provide the user with the boxed answer. Respond in the following format: <think> … </think> <answer> \boxed{…} </answer>. User: {problem}
HumanEval
System: Please reason step by step internally. Then output ONLY the final Python code that completes the task within triple backticks. Do not include explanations, markdown, or \boxed{}. User: Complete the following Python function: {prompt}
GPQA
System: You are a helpful AI Assistant. Please reason step by step, and put your final answer within \boxed{}. User: Answer the following multiple choice question. The last line of your response should be of the following format: ‘\boxed{$LETTER}’ (without quotes) where LETTER is one of ABCD (ex. ‘\boxed{A}’). Think step by step before answering. {question} A) {choice_A} B) {choice_B} C) {choice_C} D) {choice_D}
Appendix
Table 4: Task Prompt Templates
Role
Serialized Qwen prompt content
System
<|im_start|>system You are a helpful AI Assistant that provides well-reasoned and detailed responses. You first think about the reasoning process as an internal monologue and then provide the user with the boxed answer. Respond in the following format: <think> … </think> <answer> \boxed{…} </answer>. <|im_end|>
User query and positive context
<|im_start|>user Solve for x : 2x+1=32 . This is an example for a response with excellent reasoning: <|im_end|>
Assistant prefix
<|im_start|>assistant
Appendix
Table 5: Example serialization of a MATH500-style query with the positive contrastive context under the Qwen chat template.
Method
Hyperparameters
Standard sampling
N=1 , W=1 , K=32 , τ=1.0
Low-temperature sampling
N=1 , W=1 , K=32 , τ=0.25
Beam search
N=16 , W=4 , K=32 , τ=1.0 , 1/α=0
Contrastive beam search (low α )
N=16 , W=4 , K=32 , τ=1.0 , 1/α=0.25
Contrastive beam search (high α )
N=16 , W=4 , K=32 , τ=1.0 , 1/α=0.7
Power Sampling
Kt=16 , Mt=16 , B=32 (equivalent to K ), τ=0.25 (low) and τ=1.0 (high), α=4
Appendix
Table 6: Main hyperparameter settings for each method.
Figure 4: Pass@ k performance k∈{1,2,3,4} on MATH500 dataset for Power sampling (PS), standard beam search (BS) and contrastive beam search (CBS) using Deepseek-Math-7B-Instruct.
Figure 5: Best-of-16 accuracy for ten contrastive context pairs on MATH500, HumanEval and GPQA (Qwen2.5-7B), at the two deployed steering strengths. Dashed line: no-context baseline.
Content
Contrastive Context Pair
Reasoning
Positive: ‘‘This is an example for a response with excellent reasoning:’’ Negative: ‘‘This is an example for a response with wrong reasoning:’’
Completeness
Positive: ‘‘This is an example for a response with a complete treatment of the problem:’’ Negative: ‘‘This is an example for a response with an incomplete treatment of the problem:’’
Step verification
Positive: ‘‘This is an example for a response that carefully verifies each step of its reasoning:’’ Negative: ‘‘This is an example for a response that skips verification and rushes to a conclusion:’’
Reliability
Positive: ‘‘This is an example for a response that is reliable and trustworthy:’’ Negative: ‘‘This is an example for a response that is unreliable and error-prone:’’
Logical validity
Positive: ‘‘This is an example for a response with logically sound arguments:’’ Negative: ‘‘This is an example for a response with logical fallacies:’’
Reviewer judgment
Positive: ‘‘This is an example for a response that a careful reviewer would approve without hesitation:’’ Negative: ‘‘This is an example for a response that a careful reviewer would reject immediately:’’
Appendix
Table 7: Examples of 10 contrastive context pairs. Every pair shares same initial framing with “This is an example for a response …”. Each positive context is paired with its counterfactual negative.