Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.
Figures & tables
Figure 1: Overview. We compare three post-training paths—our off-policy distillation trajectories (M1), released Qwen3 off-policy-plus-on-policy distillation endpoints (M2), and the released DeepSeek-Math GRPO endpoint (M3)—on shared base model families. We evaluate them with cross-surface pass@ K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, together with strategy-diversity, multi-instance consistency, and verified memorisation probes.
Checkpoint
Post-training pipeline
Source
Qwen3-4B-Base
–
Qwen
Qwen3-4B-SFT-Math-45k-ep1
Off-policy distill
Ours
Qwen3-4B-SFT-Math-45k-ep2
Off-policy distill
Ours
Qwen3-4B-SFT-Math-45k-ep3
Off-policy distill
Ours
Qwen3-4B (Official)
Off-policy + On-policy distill
Qwen
Qwen3-8B-Base
–
Qwen
Table 1: Model roster. Nine checkpoints ( Ours ) instantiate off-policy distillation. In the Source column, Ours marks checkpoints we trained; the rest are released endpoints. Where a released pipeline is listed as off-policy distillation, that stage is the vendor’s own, not one of ours: DeepSeek-Math-7B-RL applies GRPO on top of DeepSeek’s SFT checkpoint, and the Qwen3 Official endpoints distil off-policy then on-policy.
Probe
Design
Items
Rollouts
T0
verbatim
100
4,800
T1
paraphrase ( 3× )
300
4,800
T2
numerical isomorphism ( 3× )
300
4,800
T3
translation ( 5 langs)
500
8,000
Table 2: Capability probe inventory. T0–T3 transform a shared 100 -problem source pool— 30 problems from AIME 2025, 30 from AIME 2026, and 40 from AMC 2023. Rollout counts are per evaluated checkpoint.
Figure 3: Qwen3 cross-surface pass@ K .
Figure 4: DeepSeek-Math-7B cross-surface pass@ K . Base, our off-policy distillation checkpoint, and the released GRPO endpoint are evaluated under the same pooled T0–T3 budget.
Family
Stage
all-3
any-1
fragile
Qwen3-4B
Base
26%
53%
51%
SFT-ep3
58%
75%
23%
Endpoint
84%
89%
6%
Qwen3-8B
Base
33%
55%
40%
SFT-ep3
65%
81%
20%
Endpoint
87%
90%
3%
Table 3: T1 multi-instance consistency (per source problem; 3 paraphrases, K=16 each). all-3 : all variants solved; any-1 : at least one variant solved; fragile = (any-1 − all-3) / any-1.
Family
Stage
maj
n u
n c
emb
cv len
Qwen3-4B
Base
19.6%
26.29
3.50
0.312
1.447
SFT-ep3
36.7%
8.36
1.00
0.123
0.390
Endpoint
70.1%
2.26
1.00
0.088
0.168
Qwen3-8B
Base
19.4%
27.09
1.31
0.189
1.627
SFT-ep3
48.4%
6.51
1.00
0.118
0.327
Endpoint
69.5%
2.08
1.00
0.091
0.147
Table 4: Strategy-diversity descriptors on the T0 K=48 rollout pool (§ 2.4 ). Answers are extracted via \boxed{⋅} / "Answer:" / "final answer is" patterns, then LaTeX-normalised. maj : self-consistency majority share—fraction of the 48 samples giving the most common extracted answer ( Wang et al. 2023 ) . n u : distinct extracted answers per problem. n c : single-link clusters on MiniLM-L6 response embeddings at cosine ε=0.3 . emb : mean pairwise embedding cosine distance. cv len : response-length coefficient of variation. Italics flag the DSMath Endpoint, where GRPO RL leaves maj unchanged ( −0.9 pp) while contracting the answer support nu by −66% (§ 3.3 ).
Figure 5: Content axis ( maj , nu ) across all checkpoints Answer extraction matches Table 4 .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Optimiser
AdamW, (β1,β2)=(0.9,0.999)
Peak learning rate
2×10−5
LR schedule
Cosine, min ratio 0.1
Warm-up
1% of total steps
Weight decay
0.1
Gradient clipping
1.0
Appendix
Table 5: SFT training hyperparameters.
Figure 6: Training (top row) and validation (bottom row) loss across three epochs of off-policy distillation.
Figure 7: Learning-rate schedule for the SFT runs. Left to right: Qwen3-4B-SFT-Math-45k , Qwen3-8B-SFT-Math-90k , and DSMath-7B-SFT-hybrid
Figure 8: Gradient norm for the SFT runs (clip threshold 1.0 , see Table 5 ). No divergence over three epochs.
Family
Stage
Post-training
pass@1
pass@16
pass@48
pass@96
pass@144
pass@224
ratio
probe set
T0
T0
T0
T1 ∪ T2
T0 ∪ T1 ∪ T2
∪ T3
Qwen3-4B
Base
–
10.5
39
49
62
66
68
6.48 ×
SFT-ep1
Off-policy distill
36.6
64
70
78
81
85
2.32 ×
SFT-ep2
Off-policy distill
38.9
66
75
80
81
84
2.16 ×
SFT-ep3
Off-policy distill
38.7
67
76
82
84
85
2.19 ×
Endpoint
Off + On distill
68.4
87
90
93
93
94
1.37 ×
Appendix
Table 6: Capability across three post-training mechanisms. Cross-surface pass@K pooled across T0/T1/T2/T3 surfaces (see § 3 ). The ratio column reports pass@224/pass@1 , a per-checkpoint sample-efficiency multiplier (computed from unrounded values).
Figure 9: Cross-mechanism pass@ K pooled across T0+T1+T2+T3 ( Kmax=224 )
Figure 10: Multi-instance consistency across T0, T1, T2. Each panel uses three instances of K=16 samples per source problem. T1/T2: three paraphrase/isomorphism variants. T0: the K=48 original-prompt pool partitioned into three disjoint 16 -sample batches so the per-instance budget matches T1/T2. Green: all three instances solved (at least one correct sample each); orange: some-but-not-all; grey: unsolved on all three. All three panels show the same monotone Base → SFT-ep3 consolidation, with the Qwen3 endpoints further lifting all three transforms; the DSMath RL endpoint stays within ±4 pp of SFT-ep3 on all three (T0 −4 , T1 −1 , T2 +3 pp), consistent with § 3.2 .
Figure 11: Representative T3 case.
Source problem ( amc23__000 , English T0; same item as Figure 11 ): Cities A and B are 45 miles apart. Alicia lives in A and Beth lives in B . Alicia bikes towards B at 18 miles per hour. Leaving at the same time, Beth bikes toward A at 12 miles per hour. How many miles from City A will they be when they meet? Ground truth: 27 . Model: Qwen3-4B-SFT-Math-45k-ep3 (off-policy distilled, our SFT-ep3 checkpoint). Sampled at T=0.7 , top-p=0.95 . Sample 40/48 , 406 new tokens, scored correct .
Full response (verbatim, 1,337 characters; line breaks preserved): <think>
Okay, let’s see. So, Alicia and Beth are biking towards each other from cities A and B , which are 45 miles apart. Alicia starts from A at 18 mph, and Beth starts from B at 12 mph. We need to find out how many miles from A they will meet each other. Hmm, this seems like a relative speed problem. When two objects move towards each other, their speeds add up. So, the combined speed at which they are approaching each other is 18+12=30 mph. That makes sense. Now, the total distance between them is 45 miles. If they’re moving towards each other at 30 mph, how long will it take for them to meet? Well, time is distance divided by speed. So, time =45 miles /30 mph. Let me calculate that. 45 divided by 30 is 1.5 hours. So, they will meet after 1.5 hours. Now, we need to find out how far Alicia has traveled in that time. Since speed is distance divided by time, distance = speed × time. Alicia’s speed is 18 mph, and time is 1.5 hours. Let me compute that. 18×1.5 . Hmm, 18 times 1 is 18 , and 18 times 0.5 is 9 . So, 18+9=27 miles. So, Alicia will have biked 27 miles from A when they meet. Alternatively, we could check Beth’s distance. Beth’s speed is 12 mph. 12×1.5 is 18 miles. 27+18=45 , which checks out. So that’s correct. Therefore, the answer should be 27 miles from City A .
</think> Answer: 27
Appendix
Figure 12: Representative model response with reasoning. Off-policy distilled checkpoint (Qwen3-4B SFT-ep3) on the same source problem as Figure 11 .
Model
Metric
en
stronger non-en
weaker non-en
en
zh
es
sv
ar
is
ta
ur
sw
Qwen3-4B Base
pass@1
9.6
11.9
15.0
13.5
8.4
2.1
5.9
7.6
1.8
pass@16
37
45
48
42
38
15
33
38
15
Qwen3-4B SFT-ep3
pass@1
37.3
27.8
38.4
38.4
34.0
31.7
34.6
32.9
15.4
pass@16
69
53
66
68
69
64
63
67
47
Qwen3-4B Endpoint
pass@1
69.2
64.8
68.9
70.1
69.8
58.8
58.2
58.6
31.4
Appendix
Table 9: Multilingual cross-surface reasoning, 9 models ×9 languages. Each model contributes two rows: pass@1 and pass@16 .
Figure 13: Qwen3-4B family. pass@1 (left) and pass@16 (right) across 9 languages. Bars: Base / SFT-ep3 (ours) / Endpoint (on-policy distillation).
Figure 14: Qwen3-8B family. Same layout as Figure 13 ; Base / SFT-ep3 / Endpoint comparison at 8B scale.
Figure 15: DeepSeek-Math-7B family. Endpoint is the released GRPO RL checkpoint. y -axis truncated to [0,40]% to keep sub-percent differences visible at smaller pass@1 baselines.
Figure 16: Multilingual imbalance across mechanisms. We plot pass@16 over English, stronger non-English languages (zh, es, sv, ar), and weaker non-English languages (is, ta, ur, sw) for each post-trained model. (a) Absolute scores show the family-level performance gap. (b) English-relative retention ( pass@16lang/pass@16en ) normalises across families. Weaker-language gaps persist after post-training.
Figure 17: 9-language pooled pass@ K split by difficulty subset. Each panel pools 9 languages ( 144 samples per source) and plots source-level pass@ K for K=1…144 . Rows: Qwen3-4B / Qwen3-8B / DSMath-7B; columns: AIME 2025+2026 ( n=60 ) and AMC 2023 ( n=40 ). On DSMath, GRPO RL leads at small K on AMC (sharpening) and SFT-ep3 first matches RL at K≈24 , leading at every larger K . Qwen3 endpoints stay monotone-above SFT-ep3 throughout.
Family
Epochs
Loss
MIN-K%
Zlib
ROC AUC vs. unseen
Qwen3-4B
0
0.562
0.627
0.513
10
0.603
0.648
0.548
20
0.596
0.644
0.540
Qwen3-8B
0
0.576
0.592
0.548
10
0.613
0.606
0.583
Appendix
Table 10: Forward-pass MIA ROC AUC across the overfit grid. 200 seen vs. 200 unseen items per cell, one forward pass per item (no sampling). All entries stay within [.51,.65] , below the stronger pretraining-data detection signals reported in prior work (e.g., Min-K% Prob reaching AUC 0.88 for copyrighted-book detection ( Shi et al. 2024 ) ). Qwen3 Loss AUC climbs +3 – 4 pp from 0 to 20 overfit epochs (an overfit-sensitive signal), while DSMath Loss AUC starts at ≈0.65 without overfit and stays essentially flat (baseline corpus statistics, not memorisation).
School of Computer Science and Technology, Xi’an Jiaotong University, China · Shaanxi Province Key Laboratory of Big Data Knowledge Engineering · Zhongguancun Academy, Beijing, China +1