Model Discovery Generation Digestion D i g . − D i s c . \mathrm{Dig.}-\mathrm{Disc.} Dig. − Disc. Execution E x e c . − G e n . \mathrm{Exec.}-\mathrm{Gen.} Exec. − Gen. [0pt][0pt] OpenAI gpt-5.4 82.42 82.42 82.42 67.03 67.03 67.03 100.00 100.00 100.00 (+17.58) 86.81 86.81 86.81 (+19.78) gpt-5.4-mini 59.34 59.34 59.34 50.55 50.55 50.55 97.80 97.80 97.80 (+38.46) 71.98 71.98 71.98 (+21.43) gpt-5.4-nano 41.76 41.76 41.76 44.51 44.51 44.51 97.25 97.25 97.25 (+55.49) 68.68 68.68 68.68 (+24.17) gpt-oss-20B 34.62 34.62 34.62 43.41 43.41 43.41 91.21 91.21 91.21 (+56.59) 60.99 60.99 60.99 (+17.58) [0pt][0pt] Qwen Qwen3.6-27B 24.73 24.73 24.73 52.75 52.75 52.75 92.31 92.31 92.31 (+67.58) 78.57 78.57 78.57 (+25.82) Qwen3.5-27B 28.57 28.57 28.57 47.80 47.80 47.80 93.41 93.41 93.41 (+64.84) 71.98 71.98 71.98 (+24.18) Qwen3.5-9B 13.19 13.19 13.19 38.46 38.46 38.46 79.67 79.67 79.67 (+66.48) 61.54 61.54 61.54 (+23.08) Qwen3.5-4B 6.59 6.59 6.59 26.37 26.37 26.37 70.88 70.88 70.88 (+64.29) 56.04 56.04 56.04 (+29.67) [0pt][0pt] DeepSeek-R1 Distill R1-0528-8B 6.59 6.59 6.59 17.03 17.03 17.03 68.68 68.68 68.68 (+62.09) 39.56 39.56 39.56 (+22.53) R1-Distill-32B 6.04 6.04 6.04 18.13 18.13 18.13 65.38 65.38 65.38 (+59.34) 43.41 43.41 43.41 (+25.28) R1-Distill-14B 4.40 4.40 4.40 15.93 15.93 15.93 67.03 67.03 67.03 (+62.63) 39.01 39.01 39.01 (+23.08) R1-Distill-7B 7.14 7.14 7.14 13.19 13.19 13.19 48.90 48.90 48.90 (+41.76) 32.42 32.42 32.42 (+19.23) We evaluate 12 models spanning three lineages and a broad capability range. The OpenAI lineage includes the closed-source GPT-5.4 ( OpenAI, 2026b ) , GPT-5.4-Mini ( OpenAI, 2026a ) , and GPT-5.4-Nano ( OpenAI, 2026a ) , alongside the open-weight GPT-OSS-20B ( OpenAI, 2025 ) . The Qwen lineage comprises Qwen-3.6-27B ( Qwen Team, 2026b ) and Qwen-3.5-{27B, 9B, 4B} ( Qwen Team, 2026a ) . The DeepSeek lineage features DeepSeek-R1-0528-Qwen3-8B ( DeepSeek-AI, 2025 ) and DeepSeek-R1-Distill-Qwen-{32B, 14B, 7B} ( DeepSeek-AI, 2025 ) . We use high reasoning effort when available and a maximum generation budget of 120k tokens per problem.
3.2 Results Analysis
Table 3.1 presents the comprehensive evaluation results of all models on Prim . Performance in the Discovery and Digestion dimensions is assessed by the model’s capacity to isolate the underlying mathematical primitive, utilizing the scoring protocol detailed in § 2.3 . Meanwhile, Generation and Execution are evaluated based on standard final-answer accuracy, following the HLE benchmark ( Center for AI Safety et al., 2026 ) . Further details on the diagnostic evaluation and additional analyses are provided in Appendix B .
Figure 2: Joint composition of Generation and Discovery outcomes across models. Finding 1: Answer accuracy masks distinct capability profiles. Final-answer accuracy conflates structural discovery with downstream execution. As shown in Figure 2 , Discovery and Generation are positively associated overall, yet models with similar solving accuracy can exhibit sharply different structural capabilities. For example, Qwen3.6-27B and gpt-5.4-mini achieve comparable Generation performance, while differing by over 30 30 30 percentage points in Discovery. Conversely, GPT-5.4 achieves substantially higher Discovery than Generation. The problem-level decomposition in Figure 2 reveals similarly distinct profiles across model families. The OpenAI lineage is more primitive-forward , with successful Generation frequently accompanied by correct Discovery and a substantial fraction of cases with correct Discovery but failed Generation ( D + G − D^{+}G^{-} D + G − ). In contrast, Qwen exhibits more cases with failed Discovery but successful Generation ( D − G + D^{-}G^{+} D − G + ), despite competitive Generation performance. The DeepSeek-R1 distilled models are more heavily concentrated in joint Discovery–Generation failures. Thus, similar top-line accuracy can arise from markedly different underlying capability profiles. Figure 3: Discovery versus Generation across all evaluated models on Prim . Qualitative analysis further clarifies these asymmetries. For gpt-5.4, D + G − D^{+}G^{-} D + G − cases primarily reflect two post-primitive failure modes: ❶ technical knowledge gaps , where the correct structure is identified but a required lemma, theorem condition, or domain-specific fact is missing; and ❷ procedural execution failures , such as algebraic errors, missed edge cases, or unresolved logical gaps. For Qwen3.6-27B, among D − G + D^{-}G^{+} D − G + cases, the decisive structural idea often emerges after an initial attempt without using it, typically around one quarter into the reasoning trace; in most remaining cases it appears at the outset, and only rarely is it absent entirely. The corresponding cold Discovery output often captures a partial form of the correct idea but misses its decisive component. Together with the longer reasoning traces in Figure 12 , this pattern is consistent with a more grind-first profile, where the relevant structure tends to emerge during extended procedural exploration rather than being isolated before execution. Finding 2: Correct primitives unlock latent execution capacity. To examine whether models possess downstream capabilities that are not realized during direct Generation, we evaluate Execution when the correct Mathematical Primitive is explicitly provided. As shown in Table 3.1 , primitive guidance improves performance by 17.58 17.58 17.58 to 29.67 29.67 29.67 percentage points across all 12 models. Notably, under the same primitive-guided setting, the Qwen models achieve Execution performance competitive with the closed-source gpt-5.4 variants despite substantially weaker Discovery. This gap shows that difficulty independently identifying the relevant mathematical structure can coexist with strong downstream problem-solving capability once that structure is available. More broadly, these gains indicate that many Generation failures mask downstream solving capacity that models can successfully exercise when the relevant mathematical structure is supplied. Model Generation Execution Self-Generated Primitive Teacher Plan Teacher Primitive Gold Primitive gpt-5.4-mini 50.55 50.55 50.55 48.90 48.90 48.90 (-1.65) 60.44 60.44 60.44 (+9.89) 63.74 63.74 63.74 (+13.19) 71.98 71.98 71.98 (+21.43) gpt-5.4-nano 44.51 44.51 44.51 46.15 46.15 46.15 (+1.64) 54.40 54.40 54.40 (+9.89) 60.99 60.99 60.99 (+16.48) 68.68 68.68 68.68 (+24.17) Qwen3.6-27B 52.75 52.75 52.75 41.21 41.21 41.21 (-11.54) 56.04 56.04 56.04 (+3.29) 64.84 64.84 64.84 (+12.09) 78.57 78.57 78.57 (+25.82) Table 2: Generation and Execution performance on Prim under different forms of reasoning support. Detailed results are presented in Table B.3.3 in Appendix B.3 . We further investigate whether this rescue can be attributed simply to task decomposition or generic procedural scaffolding. Table 2 compares four intervention settings. Conditioning the model on its own generated primitive provides little benefit: it yields only a marginal gain for gpt-5.4-nano and substantially degrades Qwen3.6-27B, indicating that a forced discover-then-execute decomposition alone is insufficient. We then compare two forms of external guidance generated by the same teacher model (gpt-5.4) from the problem alone: a step-by-step Teacher Plan , which serves as a procedural hint by specifying the major solution steps, and a Teacher Primitive , which is generated according to the definition of primitive. Although the Teacher Plan improves over direct Generation, the Teacher Primitive yields substantially larger gains across the evaluated models. The Gold Primitive further maximizes this effect, pushing accuracy to 71.98 % 71.98\% 71.98% , 68.68 % 68.68\% 68.68% , and 78.57 % 78.57\% 78.57% for the representative models. This contrast provides crucial empirical validation for our premise in § 2.1 : a mathematical primitive is fundamentally distinct from, and vastly more effective than, a generic plan or hint. The rescue effect therefore cannot be attributed to task decomposition or the procedural scaffold evaluated here alone; exposing the load-bearing mathematical structure provides a substantially stronger signal for downstream reasoning. Finding 3: Independent structural discovery is the dominant bottleneck. As shown in Table 3.1 , models perform substantially better on Digestion , recovering the Mathematical Primitive from a correct solution, than on Discovery , where the same primitive must be identified from the problem alone. This asymmetry is particularly pronounced outside the strongest OpenAI models. Even models with weak Discovery can often recover the organizing primitive once the relevant reasoning is exposed through a completed solution; for instance, Qwen3.6-27B rises from 24.73% Discovery to 92.31% Digestion and R1-Distill-32B from 6.04% to 65.38%.
Figure 4: Decomposition of Generation failures by Discovery ( D D D ) and Execution ( E E E ). Detailed results can be found in Table B.3.3 in Appendix B.3 . These results reveal a pronounced gap between retrospective recognition and prospective discovery of mathematical structure: models can often recognize the key idea once it is instantiated in a valid solution, yet struggle to identify it independently before the derivation is known.
To further localize this bottleneck, we decompose Generation failures according to the joint outcomes of Discovery ( D D D ) and Execution ( E E E ), as shown in Figure 4 . A striking 83.6 % 83.6\% 83.6% of all failures fall into the D − D^{-} D − regime, indicating that failed primitive discovery dominates the failure landscape. Importantly, a substantial fraction of these cases are discovery-limited failures ( D − E + D^{-}E^{+} D − E + ): the model fails to identify the correct primitive independently, yet succeeds once it is supplied. This reveals that, for many failures, latent execution capacity is already present but remains inaccessible because of the upstream bottleneck in primitive discovery. 3.3 Further Analysis: Which Failures Are Repairable?
While our preceding diagnosis reveals the failure modes of LLMs during inference, a natural follow-up question arises: how does post-training impact these mathematical reasoning failures? Of particular interest are discovery-limited failures, which exhibit substantial room for improvement (§ 3.2 ). We hypothesize that these failures lie just at the boundary of the models’ reasoning capabilities, making them highly amenable to repair through post-training. Therefore, we utilize Supervised Fine-Tuning (SFT) and On-Policy Self-Distillation (OPSD) ( Zhao et al., 2026 ) to investigate how different failure modes uniquely respond to these post-training paradigms.
Setup. Our training corpus contains 709 Mathematics Ph.D. qualifying-examination problems from 1991–2026, each paired with a human-authored proof (median length: 118 words); additional details are provided in Appendix C . We use qualifying-examination problems because their open-ended, proof-oriented nature closely matches the difficulty and structural reasoning demands of Prim . Since such problems generally lack algorithmically verifiable short answers, we focus on text-supervised post-training rather than outcome-verified methods such as RLVR. We evaluate SFT and OPSD on Qwen-3.5-27B, 9B, 4B ( Qwen Team, 2026a ) , using the default OPSD configuration, bfloat16 precision, a fixed seed of 42 42 42 , and NVIDIA RTX 6000 Ada GPUs. On-Policy Self-Distillation. Given training examples { x , y } \{x,y\} { x , y } , where x x x denotes the input problem and y y y its reference solution, both the teacher and student policies are initialized from the same base model π 0 \pi_{0} π 0 . At each decoding step t t t , the student autoregressively generates an on-policy response, inducing the token distribution P t : = π S ( ⋅ ∣ y ^ < t , x ) P_{t}:=\pi_{S}(\cdot\mid\hat{y}_{<t},x) P t := π S ( ⋅ ∣ y ^ < t , x ) , while the teacher distribution Q t : = π T ( ⋅ ∣ y ^ < t , x , y ) Q_{t}:=\pi_{T}(\cdot\mid\hat{y}_{<t},x,y) Q t := π T ( ⋅ ∣ y ^ < t , x , y ) additionally conditions on the reference solution. Notably, both policies are conditioned on the same student prefix y ^ < t \hat{y}_{<t} y ^ < t , allowing the teacher to provide token-level supervision along the student’s own reasoning trajectory. OPSD minimizes the divergence between these distributions: L O P S D = 1 T ∑ t = 1 T D ( P t , Q t ) = 1 T ∑ t = 1 T D ( π S ( ⋅ ∣ y ^ < t , x ) , π T ( ⋅ ∣ y ^ < t , x , y ) ) , \displaystyle\mathcal{L}_{\mathrm{OPSD}}=\frac{1}{T}\sum_{t=1}^{T}D(P_{t},Q_{t})=\frac{1}{T}\sum_{t=1}^{T}D\left(\pi_{S}(\cdot\mid\hat{y}_{<t},x),\pi_{T}(\cdot\mid\hat{y}_{<t},x,y)\right), L OPSD = T 1 t = 1 ∑ T D ( P t , Q t ) = T 1 t = 1 ∑ T D ( π S ( ⋅ ∣ y ^ < t , x ) , π T ( ⋅ ∣ y ^ < t , x , y ) ) , where D D D is instantiated as Kullback–Leibler (KL), reverse KL (rKL), or Jensen–Shannon (JSD). Specifically, KL uses D K L ( Q t ∥ P t ) D_{\mathrm{KL}}(Q_{t}\parallel P_{t}) D KL ( Q t ∥ P t ) , rKL uses D K L ( P t ∥ Q t ) D_{\mathrm{KL}}(P_{t}\parallel Q_{t}) D KL ( P t ∥ Q t ) , and JSD is symmetric in P t P_{t} P t and Q t Q_{t} Q t . Results. Across the 27B, 9B, and 4B models, the base models leave 341 341 341 failure instances in Prim . We focus on the 294 294 294 cases with failed Discovery: 147 147 147 D − E + D^{-}E^{+} D − E + and 147 147 147 D − E − D^{-}E^{-} D − E − cases. We define repair as an incorrect baseline Generation response becoming correct after post-training. As shown in Table 3 , SFT and OPSD repair 20.4 % 20.4\% 20.4% and 21.1 % 21.1\% 21.1% of D − E + D^{-}E^{+} D − E + cases, roughly three Failure type SFT OPSD D − E + D^{-}E^{+} D − E + 20.4% 21.1% D − E − D^{-}E^{-} D − E − 6.8% 8.2% Table 3: Post-training repair rates of SFT and OPSD across initial Prim failure quadrants. times their respective rates of 6.8 % 6.8\% 6.8% and 8.2 % 8.2\% 8.2% for D − E − D^{-}E^{-} D − E − cases. This ordering holds at each model scale, identifying discovery-limited failures as a promising target for post-training. However, net gains in Generation remain modest (as shown in Table 4 ) because some initially correct cases become incorrect, offsetting part of these repairs. These regressions are consistent with post-training drift and suggest that supervision from reference proofs should be transferred more selectively. 4 Internalizing Primitive-Guided Reasoning
In this section, we introduce Absorb , a novel variant of on-policy self-distillation that improves the mathematical problem-solving capability of LLMs by internalizing primitive-guided reasoning, motivated by the observations in § 3 . Extensive experiments with Qwen models as backbones show that Absorb consistently outperforms strong baselines on a diverse set of mathematical reasoning benchmarks (§ 4.4 ). Further implementation details and ablation results are provided in Appendix E .