A helps B while B hurts A: directed transfer in instruction-tuning mixture
Organizations: Juna.ai, Kastanienallee 32, 10435 Berlin, Germany · Computational Systems Neuroscience, Institute of Zoology, University of Cologne, Germany
Abstract
Adapting a language model to a specialized corpus means choosing which instruction-tuning tasks to train on under a fixed budget, and testing one choice costs a fine-tuning run. Common heuristics add more source tasks or pick sources similar to the target. The first assumes transfer is never negative; the second, that it is symmetric. We show that both assumptions fail: task can help task while hurts , so helpfulness is a signed property of ordered source--target pairs. We introduce the transfer map, a signed estimate of how much each source helps or hurts each held-out target. We fit the map in hundreds of fine-tuning runs on Qwen3 and Mistral models from 0.6B to 32B parameters, with all sources drawn from one corpus and no training examples from the target. The map predicts a held-out target's accuracy on unseen mixtures: recorded before those runs, its predictions have less than half the error of a mixture-agnostic baseline. The map is specific to its target and corpus but transfers across model scale: a mixture selected in advance at one size beats training on all source tasks at every other size we tested. Transfer is thus a property of the data. The map selects the tasks that help and drops the one that interferes: accuracy on the reasoning targets (causal explanation, multi-hop questions and methodological critique) rises by up to 14 percentage points over training on all source tasks.
Figures & tables
| Condition | Target | Interferer | Uniform | Single-best | Skill-It | TaskShop | Embed-sim |
|---|---|---|---|---|---|---|---|
| Qwen3-32B | why | mc-mc | |||||
| Qwen3-32B | mhop | err | |||||
| Qwen3-32B | err | mhop | |||||
| Mistral-24B | why | err | |||||
| Mistral-24B | mhop | err | |||||
| Mistral-24B | err | mhop |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Predictor | MAE (pp) | Pearson | Origin slope |
|---|---|---|---|
| Recorded map prediction | |||
| Zero effect | — | — | |
| Best constant, fitted post hoc | — | ||
| Count, linear | |||
| Count, quadratic |
| Scale: agreement with 32B map | Transport: between conditions | ||
| Qwen3-0.6B | Qwen Mistral | ||
| Qwen3-1.7B | Qwen PubMed | ||
| Qwen3-4B | Mistral PubMed | ||
| Qwen3-8B | |||
| Out-of-fold gain over the widest mixture of the left-out design (pp) | |||
| Qwen3-32B, full probe | Qwen3-32B, runs | ||
| Transported B selection | Refit at this size | ||||||
|---|---|---|---|---|---|---|---|
| Size | Target | Recorded | Realized | 95% CI | Realized | 95% CI | |
| 0.6B | why | ||||||
| 0.6B | mhop | ||||||
| 0.6B | err | ||||||
| 1.7B | why | ||||||
| 1.7B | mhop | ||||||
| Condition | Target | Sources kept | Recorded | Realized | 95% CI | |
|---|---|---|---|---|---|---|
| Qwen3-32B | why | clz , def , err , mhop | ||||
| Qwen3-32B | mhop | clz , def , match , why | ||||
| Qwen3-32B | err | def , match , sc-mc , why | ||||
| Mistral-24B | why | clz , def , mc-mc , mhop , sc-mc | ||||
| Mistral-24B | mhop | clz , def , mc-mc | ||||
| Mistral-24B | err | all but mhop |
| Qwen3-32B | Mistral-24B | PubMed | ||||
|---|---|---|---|---|---|---|
| Model of the mixture | MAE | MAE | MAE | |||
| Size, linear | ||||||
| Size, quadratic | ||||||
| Size, categorical | ||||||
| Source presence | ||||||
| Presence share | ||||||
| Condition | Target | Accuracy at / … / | Slope in | Diff. | |
|---|---|---|---|---|---|
| Qwen3-32B | why | / / / / / | |||
| Qwen3-32B | mhop | / / / / / | |||
| Qwen3-32B | err | / / / / / | |||
| Qwen3-32B | def | / / / / / | |||
| Mistral-24B | why | / / / / / | |||
| Mistral-24B | mhop | / / / / / |
| Opposite signs | ||||||
|---|---|---|---|---|---|---|
| Condition | Skew | Directional | Estimates | Diff. CI | Both CIs | |
| Qwen3-32B | ||||||
| Mistral-24B | ||||||
| PubMed | ||||||
| Pair | First | 95% CI | Second | 95% CI | Difference | 95% CI |
|---|---|---|---|---|---|---|
| why , mhop | ||||||
| mhop , def ∗ | ||||||
| err , mhop | ||||||
| err , why ∗ | ||||||
| why , def ∗ | ||||||
| why , mc-mc |
| Cond. | Target | Competitor | Interferer in mix | target | other-6 | |
|---|---|---|---|---|---|---|
| Qwen | why | uniform | yes ( mc-mc ) | |||
| Qwen | why | single_best | no | |||
| Qwen | why | skillit | no | |||
| Qwen | why | taskshop_k3 | no | |||
| Qwen | why | taskshop | no | |||
| Qwen | why | taskshop_k5 | no |
| Pair | Size | Sources | Target | Predicted | Realized [95% CI] | Last vs shuffled [95% CI] | Status |
| P1 at three model sizes | |||||||
| P1 | 4B | def + err | mhop | confirmed | |||
| P1r17 | 1.7B | def + err | mhop | confirmed | |||
| P1r8 | 8B | def + err | mhop | confirmed | |||
| Other recorded pairs | |||||||
| P2 | 4B | mhop + sc-mc | why | covers | |||
| Task | Capability | Items | Ref. words | Scoring |
|---|---|---|---|---|
| def | definition with ontology fields | Eq. 1 | ||
| clz | fill-in-the-blank span | normalized exact match | ||
| sc-mc | one correct option | option match | ||
| mc-mc | multiple correct options | option-set match | ||
| match | term–definition pairing | pair-set match | ||
| mhop | two-evidence reasoning | Eq. 1 |