Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and lucky guesses. We propose Collaborative Reasoning Distillation (CRD), a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other's reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths. Students are trained via Reasoning Quality Optimization (RQO) with budget constraints. Our model, CRD-4B, achieves 97.3% on MATH-500 and 70.3% on AIME'25, surpassing baselines while using only 50K training examples, up to 12 times smaller than the datasets of comparable models.
Figures & tables
Figure 1 : Overview of the CRD Framework. Multiple teachers iteratively critique and refine each other’s reasoning through Cross-Feedback . An LLM-as-a-Judge then assigns step-wise quality scores, and the Thinking Path algorithm stitches high-quality steps across teachers while pruning flawed steps. The curated dataset is used to train students via RQO with Optimal Budget curriculum.
Math
Code
Graduate-Level STEM
Model
MATH-500
AIME’24
AIME’25
LCB
GPQA-Diamond
Tiny Scale ( < 2B)
DeepSeek-R1-Distill-7B
92.8
55.5
–
37.6
49.1
Gemini-Flash-Distill-1.5B
85.3
32.5
24.6
17.2
31.6
DeepSeek-R1-Distill-1.5B
83.9
28.9
22.0
16.9
32.9
ProRL-1.5B †
91.9
48.1
33.3
23.8
41.7
Table 1 : Performance comparison across benchmarks (Pass@1). We report accuracy (%) on mathematical, coding, and general reasoning benchmarks. The best results are highlighted in bold . Gray rows indicate DeepSeek-R1 and distilled models included as reference points for parameter efficiency comparison.
Table 3
Figure 2 : ( a ) CRD achieves comparable performance to DeepSeek-R1-Distill models with 4.7–8 × fewer parameters. CRD-1.5B approaches the 7B model on MATH-500 (4.7 × ), while CRD-4B surpasses the 32B model (8 × ). ( b ) CRD achieves superior performance with 12 × less data than DeepSeek-R1-Distill (600K reasoning samples) and 2.7 × less than ProRL (136K), while outperforming both on AIME’25 and LCB.
BoxNet
pass@1
pass@2
pass@4
DeepSeek-R1-Distill-1.5B
0.0
0.0
0.0
ProRL-1.5B
7.1
13.8
22.4
CRD-1.5B
6.8
14.6
25.3
Table 4 : Generalization beyond the training distribution. Left: BoxNet planning (100 procedurally generated instances, our re-evaluation of all models, temperature 0.6, top- p 0.95, 32K tokens). Right: HLE text-only subset (official protocol).
Method
MATH-500
AIME’25
LCB
GPQA-Diamond
Naive SFT
90.5
33.4
25.7
38.9
RQO (CRD)
91.1
35.1
27.2
39.7
Table 5 : Same-data comparison of training objectives (DeepSeek-R1-Distill-Qwen-1.5B, 50K Thinking Path trajectories).
Figure 7
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Score
Grade
Description
Decision Boundary
10.0
Rigorous
Flawless execution, all constraints satisfied, all cases covered.
N/A (Perfect)
8.0–9.9
Valid
Correct logic with minor inefficiency (redundancy, non-standard notation).
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China · 2The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China