cs.LGSep 29, 2026

Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning

Authors: Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu, Said Mammar, Pascal Bouvry

Organizations: University of Luxembourg · Seafill Open-Source Community · Université Paris-Saclay

Abstract

Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

    Oct 1, 2026Shuo Xing, Zilin Dai, Chengyuan Qian +7Mathematical ReasoningLLM Reasoning Strategies

  2. From Rollouts to Recipes: Self-Contained Post-Training for LLMs

    Sep 1, 2026Yifei Li, Lingling Zhang, Muye Huang +3Post-TrainingReinforcement Learning Post-Training