How code helps different tasks? A decompositional lens on LLM post-training
Organizations: Sun Yat-sen University · Shenzhen Loop Area Institute · The Chinese University of Hong Kong, Shenzhen · Shenzhen Research Institute of Big Data
Abstract
Evaluating code data as a single corpus can obscure which types of code data benefit which models and downstream tasks. Effective data selection requires understanding both the benefits of individual categories and whether these benefits persist when categories are combined. We introduce a decompositional lens for studying these effects in LLM post-training. We first decompose an execution-verified code corpus into interpretable categories based on the computational patterns of its solutions. Through controlled fine-tuning experiments, we compare individual categories with a balanced mixture across instruction-tuned models on question answering, mathematics, and code generation. The resulting response maps reveal recurring gains in average question-answering performance, while the same category can improve one model or task and degrade another. The best-performing category also varies with the starting model and target task. We then compose compact mixtures guided by these results and examine whether benefits observed in individual categories persist under joint training. On selected model--task pairs, mixtures whose constituents each improve the target task outperform both their best constituent and full-corpus training while using roughly 10--15% of the full corpus. These exploratory findings illustrate a \emph{less is more} pattern and highlight how the value of code data in post training depends on which categories are combined for which model and task.
Figures & tables
| Category | Representative patterns | |
|---|---|---|
| Math / NT | Modular arith. Counting Transforms | 2,075 |
| Dynamic prog. | Recurrence Memoization Tables | 3,562 |
| Graph / stateful | Traversal Connectivity Paths States | 2,775 |
| Greedy / search | Greedy Binary/local search Rules | 1,706 |
| Impl. / utilities | Direct impl. Bookkeeping I/O glue | 5,621 |
| Hashing / sets | Hash maps Counting Sets Dedup | 5,077 |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Recorded value |
|---|---|
| Source and labels | 35,974 cleaned KodCode training examples; selection by primary category only. |
| Single-category configurations | 1,706 examples each; sampling without replacement. |
| Balanced mixture | 1,706 examples drawn from the single-category subsets: six quotas of 171 and four of 170. |
| Sampling seed | 20260730; no auxiliary-label reselection. |
| Input and loss | Native model chat template; assistant-only loss. |
| LoRA | Rank 16, scaling parameter 32, dropout 0.05; attention query, key, value, and output projections. |
| Configuration | Scored | Metric | Scoring mode | Max. tokens |
|---|---|---|---|---|
| ARC-Challenge | 1,172 | Accuracy | Option log probability | — |
| FinQA | 1,133 | Accuracy | Generated answer | 128 |
| HumanEval | 164 | Pass@1 | Generated code | 2,048 |
| LegalBench | 1,689 | Accuracy | Option log probability | — |
| MATH-M | 105 | Accuracy | Generated answer | 512 |
| MBPP+ | 378 | Pass@1 | Generated code | 2,048 |
| (a) QA (5 configurations) | ||||||
| Training recipe | Qwen2.5-7B | Qwen3-4B | Llama3.1-8B | Gemma2-9B | Llama3.2-3B | Qwen1.5-7B |
| Base | 43.3 | 43.1 | 44.8 | 44.5 | 36.2 | 36.2 |
| Balanced mix | 46.2 | 45.4 | 45.4 | 46.7 | 37.7 | 36.9 |
| Math / NT | 46.2 | 45.9 | 45.2 | 47.0 | 37.2 | 36.9 |
| Dynamic programming | 46.0 | 45.4 | 45.4 | 46.7 | 37.6 | 36.4 |
| Graph / stateful | 45.9 | 44.9 | 45.3 | 46.8 | 37.5 | 36.5 |
| Arm | Examples | Updates | Composition |
|---|---|---|---|
| Base | 0 | 0 | Original checkpoint |
| Personalized Top3 | 5,118 | 320 | Three model-specific categories, 1,706 each |
| Shared Top3 | 5,118 | 320 | Impl./util., Hash/sets, Sort/order; 1,706 each |
| Balanced | 5,118 | 320 | Math and Seq: 511 each; other eight: 512 each |
| Proportional | 5,118 | 320 | Stratified sample preserving pool proportions |
| Complete | 35,974 | 2,249 | Full source pool, one epoch |
| Model | Base | Personal. | Shared | Balanced | Proport. | Complete |
|---|---|---|---|---|---|---|
| Qwen2.5-7B | 55.782 | 58.571 | 59.165 | 58.825 | 59.035 | 56.781 |
| Qwen3-4B | 55.242 | 56.872 | 56.361 | 57.361 | 56.651 | 57.975 |
| Qwen1.5-7B | 37.145 | 37.347 | 38.407 | 38.244 | 38.066 | 37.667 |
| Llama-3.1-8B | 51.496 | 52.077 | 51.726 | 52.363 | 51.589 | 52.900 |
| Llama-3.2-3B | 44.898 | 43.971 | 44.264 | 44.209 | 43.775 | 41.373 |
| Gemma-2-9B | 53.165 | 53.988 | 53.400 | 53.438 | 53.660 | 50.935 |
| Model | Personal. Bal. | Personal. Prop. | Shared Bal. | Shared Prop. |
|---|---|---|---|---|
| Qwen2.5-7B | ||||
| Qwen3-4B | ||||
| Qwen1.5-7B | ||||
| Llama-3.1-8B | ||||
| Llama-3.2-3B | ||||
| Gemma-2-9B |
| Model | Base, 256 | Base, 1,024 | Change (pp) | Positive arms: 256 1,024 |
|---|---|---|---|---|
| Qwen2.5-7B | 76.346 | 91.054 | ||
| Qwen3-4B | 76.649 | 93.025 | ||
| Qwen1.5-7B | 59.591 | 63.154 | ||
| Llama-3.1-8B | 72.176 | 78.317 | ||
| Llama-3.2-3B | 72.328 | 76.801 | ||
| Gemma-2-9B | 84.685 | 85.444 |
| Model | Personal. | Shared | Balanced | Proport. | Complete |
|---|---|---|---|---|---|
| Qwen2.5 | |||||
| Qwen3 | |||||
| Qwen1.5 | |||||
| Llama-3.1 | |||||
| Llama-3.2 | |||||
| Gemma-2 |
| Backbone | Task | Recipe | Best single | Mixture | Complete | Base |
|---|---|---|---|---|---|---|
| Qwen2.5 | ARC-C | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.5586,0.3555,0.6523}\faIcon{font}}\penalty\hskip 2.20001ptStr}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash} | 59.64 | 61.52 | 59.13 | |
| ARC-C | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash}\,+\,\text{{\color[rgb]{0.082,0.457,0.4141}\faIcon{sort-amount-down}}\penalty\hskip 2.20001ptSort} | 59.64 | 61.18 | 59.13 | ||
| ScienceQA | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.5586,0.3555,0.6523}\faIcon{font}}\penalty\hskip 2.20001ptStr}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash} | 70.59 | 71.45 | 70.64 | ||
| ScienceQA | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash}\,+\,\text{{\color[rgb]{0.082,0.457,0.4141}\faIcon{sort-amount-down}}\penalty\hskip 2.20001ptSort} | 70.59 | 71.18 | 70.64 | ||
| GSM8K † | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash}\,+\,\text{{\color[rgb]{0.082,0.457,0.4141}\faIcon{sort-amount-down}}\penalty\hskip 2.20001ptSort} | 84.08 | 86.28 | 83.85 | ||
| MATH-H | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash}\,+\,\text{{\color[rgb]{0.082,0.457,0.4141}\faIcon{sort-amount-down}}\penalty\hskip 2.20001ptSort} | 30.15 | 34.73 | 25.19 |
| Backbone | Task | Recipe | Best single | Mixture | Complete | Base |
|---|---|---|---|---|---|---|
| Qwen2.5 | ARC-C | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.5586,0.3555,0.6523}\faIcon{font}}\penalty\hskip 2.20001ptStr} | 59.64 | 60.75 | 59.13 | |
| ScienceQA | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.5586,0.3555,0.6523}\faIcon{font}}\penalty\hskip 2.20001ptStr} | 70.59 | 70.77 | 70.64 | ||
| GSM8K † | \text{{\color[rgb]{0.4844,0.5781,0.7539}\faIcon{cogs}}\penalty\hskip 2.20001ptImp}\,+\,\text{{\color[rgb]{0.5586,0.3555,0.6523}\faIcon{font}}\penalty\hskip 2.20001ptStr} | 84.00 | 84.84 | 83.85 | ||
| Qwen3 | LegalBench | \text{{\color[rgb]{0.7891,0.4336,0.1211}\faIcon{exchange-alt}}\penalty\hskip 2.20001ptSeq}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash} | 84.67 | 85.08 | 84.25 | |
| MBPP-Simple | \text{{\color[rgb]{0.7891,0.4336,0.1211}\faIcon{exchange-alt}}\penalty\hskip 2.20001ptSeq}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash} | 78.60 | 79.38 | 77.82 | ||
| Qwen1.5 | ARC-C | \text{{\color[rgb]{0.7891,0.4336,0.1211}\faIcon{calculator}}\penalty\hskip 2.20001ptMath}\,+\,\text{{\color[rgb]{0.3672,0.6133,0.5508}\faIcon{hashtag}}\penalty\hskip 2.20001ptHash} | 57.17 | 57.76 | 54.69 |
| ARC-Challenge | ScienceQA | |||
|---|---|---|---|---|
| Backbone | Imp Base | Imp Other nine | Imp Base | Imp Other nine |
| Qwen2.5-7B | ||||
| Qwen3-4B | ||||
| Llama-3.1-8B | ||||
| Gemma-2-9B | ||||
| Llama-3.2-3B | ||||
| Backbone / item | Question and Base Imp answer | Interpretation and comparison |
|---|---|---|
| Llama-3.1 / ARC arc_00372 | One hour at 80 km/h, then one hour at 100 km/h: km/h. | Corrected equal-duration average speed; all other nine arms and the balanced mixture are wrong. |
| Qwen2.5 / ARC arc_00564 | A substance expands and then melts when heated: initial state liquid solid. | Corrected initial-state inference; all other nine arms and the balanced mixture are wrong. |
| Llama-3.2 / ScienceQA scienceqa_02187 | Given dominant , recessive , and genotype : brown red eyes. | Corrected rule application; all other nine arms and the balanced mixture are wrong. |
| Qwen2.5 / ARC arc_00821 | Imp correctly selects that plants existed before coal and oil formed; Base is wrong. | Factual improvement; all other nine single-category training configurations are wrong. |
| Llama-3.1 / ScienceQA scienceqa_02144 | Both journeys take five hours: 75 miles 60 miles as the faster journey. | Relational regression: Base is correct and Imp is wrong. |
| Qwen2.5 / ScienceQA scienceqa_01144 | Identify an object that is not a mineral: skull quartz. | Classification regression: Base and all other nine arms are correct. |
| Benchmark | R/F | R gain | F gain | R F | 95% interval |
|---|---|---|---|---|---|
| ARC-Challenge | 26/41 | ||||
| ScienceQA | 20/15 |
| Skill group | Imp Base | Imp Other nine | |
|---|---|---|---|
| Physical relations (e.g., speed, force, energy) | 141 | ||
| Experimental constraints | 61 | ||
| Genotype, phenotype, and dominance rules | 105 | ||
| Inherited/acquired trait evidence | 113 | ||
| Remaining natural-science questions | 623 |
| Positive arms | Mean change (pp) | |||
|---|---|---|---|---|
| Backbone | QA5 | QA3 | QA5 | QA3 |
| Qwen2.5-7B | 10/10 | 10/10 | ||
| Qwen3-4B | 10/10 | 10/10 | ||
| Llama-3.1-8B | 10/10 | 5/10 | ||
| Gemma-2-9B | 10/10 | 10/10 | ||
| Llama-3.2-3B | 10/10 | 10/10 | ||
| Backbone | Mean (pp) | Positive/negative | Base tokens | Trained tokens |
|---|---|---|---|---|
| Qwen2.5-7B | 10/0 | 193.3 | 148.7 | |
| Qwen3-4B | 10/0 | 181.5 | 168.6 | |
| Llama-3.1-8B | 10/0 | 177.1 | 170.6 | |
| Gemma-2-9B | 0/10 | 100.0 | 89.0 | |
| Llama-3.2-3B | 0/10 | 175.6 | 160.9 | |
| Qwen1.5-7B | 0/10 | 169.4 | 113.7 |
| Backbone | CW | WC | Assertion | Helper | Indentation | Other |
|---|---|---|---|---|---|---|
| Qwen3-4B | 154 | 36 | 131 | 10 | 0 | 13 |
| Gemma-2-9B | 85 | 129 | 53 | 5 | 27 | 0 |
| Backbone / arm / item | Observed change and interpretation |
|---|---|
| Qwen1.5 / String gsm8k_00972 | Base confuses 40 oranges kept with the number sold; the trained response correctly gives 120 sold. Both answers are complete, and the correct response is shorter (117 105 visible tokens). |
| Gemma-2 / Sorting gsm8k_01070 | The trained response halves the quantity of discounted butter twice, changing the correct answer 18 to 15. The response becomes longer (87 112 tokens). |
| Qwen3 / Imp HumanEval/26 | Base removes every value occurring more than once; the trained code instead retains one copy of each value. This changes required semantics. The same regression occurs in all ten categories and the mixture on this item. |
| Gemma-2 / Sequence HumanEval/159 | A shorter program correctly updates remaining inventory by subtracting the available amount sold. This is a successful relation update rather than a length-induced failure. |
| Gemma-2 / Sorting HumanEval/50 | The decoding formula is retained, but the prompt-provided helper is absent from the assembled test program. The recorded regression is compatible with a wrapper failure. |