Task Vector Descent: Learning from Non-IID Batches
Organizations: ETH Zurich · Technical University of Munich
Abstract
A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions that are encountered unevenly over time. In such settings, successive minibatches are temporally clustered by distribution instead of being sampled i.i.d. from the overall data mixture. Training on temporally clustered data induces a stability-plasticity tradeoff. Adapting the model to the active distribution can improve the model on the active distribution but may lead to a performance degradation on data it previously trained on. We find that this tradeoff intensifies with longer exposure to the same distribution. We therefore ask if the parameter displacement produced by such a sequence (the task vector) should be fully retained or applied partially. We compare applying the full displacement () with partial integration, which scales the task vector by before applying it to the continuing model and scales the optimizer state by the same coefficient. Across continual pretraining, pretraining from random initialization, supervised post-training, and reinforcement post-training, we find that intermediate values of often improve average continuing-model performance relative to full integration, particularly after longer same-distribution sequences. In continual-pretraining experiments with both controlled streams and naturally defined adaptation sequences, task-vector scaling outperforms full integration at the matched learning rate, showing that its benefits are not reproduced by learning-rate scaling alone.
Figures & tables
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Model | # Seeds | Primary score | |||
|---|---|---|---|---|---|---|
| Continual PT | GPT-2 | 10 | 3 | BPT gain | ||
| Persona SDPO | Qwen3-4B | 6 | 3 | Macro win rate | ||
| Reasoning SFT | Qwen2.5-3B | 4 | 5 | Best4 | ||
| Reasoning GRPO | Qwen2.5-3B | 4 | 3 | Best4 | ||
| Scratch PT | GPT-2 | 10 | 3 | BPT gain | ||
| Book-level PT | SmolLM-360M | 200 | – | 3 | BPT gain |
| Asset / checkpoint | License / terms |
|---|---|
| Models and services | |
| GPT-2 (124M) ( Radford et al., 2019 ) | Modified MIT |
| Qwen3-4B , Qwen3-8B ( Yang et al., 2025 ) | Apache 2.0 |
| Qwen2.5-1.5B , Qwen2.5-7B ( Yang et al., 2024 ) | Apache 2.0 |
| Qwen2.5-3B , Qwen2.5-3B-Instruct ( Yang et al., 2024 ) | Qwen Research License |
| Qwen2.5-Coder-1.5B , -7B ( Hui et al., 2024 ) | Apache 2.0 |
| Seed | Full-book tokens (k) | Training targets (M) | Steps per book | Total steps | ||
|---|---|---|---|---|---|---|
| Median | P90 | Median | Range | |||
| 0 | 80.9 | 221.4 | 18.49 | 9 | 1–91 | 2,353 |
| 1 | 81.1 | 219.0 | 18.98 | 9 | 1–57 | 2,410 |
| 2 | 83.2 | 242.1 | 20.41 | 9.5 | 1–158 | 2,585 |
| Category | Seed 0 | Seed 1 | Seed 2 |
|---|---|---|---|
| Novels | 32.5 | 30.0 | 32.0 |
| British Literature | 20.0 | 18.0 | 19.5 |
| American Literature | 10.0 | 10.0 | 14.0 |
| History - Modern (1750+) | 10.5 | 10.5 | 13.0 |
| Adventure | 11.5 | 12.0 | 9.5 |
| Biographies | 9.0 | 8.5 | 12.5 |
| Setting | GPUs/run | Timed runs | Hours/run | GPU-hours |
|---|---|---|---|---|
| Continual pretraining | 4 | 72 | 0.11–0.55 | 79.85 |
| Persona SDPO | 2 | 84 | 0.91–2.94 | 292.57 |
| Reasoning SFT | 4 | 125 | 0.29–3.58 | 540.36 |
| Reasoning GRPO | 4 | 39 | 4.53–8.51 | 896.76 |
| Pretraining from scratch | 4 | 78 | 1.94–2.59 | 663.71 |
| PG-19 continual pretraining | 4 | 72 | 0.49–1.09 | 202.87 |