cs.LGOct 4, 2026

Task Vector Descent: Learning from Non-IID Batches

Authors: Anton Baumann, Jonas Hübotter, Zeynep Akata, Andreas Krause

Organizations: ETH Zurich · Technical University of Munich

Abstract

A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions that are encountered unevenly over time. In such settings, successive minibatches are temporally clustered by distribution instead of being sampled i.i.d. from the overall data mixture. Training on temporally clustered data induces a stability-plasticity tradeoff. Adapting the model to the active distribution can improve the model on the active distribution but may lead to a performance degradation on data it previously trained on. We find that this tradeoff intensifies with longer exposure to the same distribution. We therefore ask if the parameter displacement produced by such a sequence (the task vector) should be fully retained or applied partially. We compare applying the full displacement (λ=1λ=1) with partial integration, which scales the task vector by λλ before applying it to the continuing model and scales the optimizer state by the same coefficient. Across continual pretraining, pretraining from random initialization, supervised post-training, and reinforcement post-training, we find that intermediate values of λλ often improve average continuing-model performance relative to full integration, particularly after longer same-distribution sequences. In continual-pretraining experiments with both controlled streams and naturally defined adaptation sequences, task-vector scaling outperforms full integration at the matched learning rate, showing that its benefits are not reproduced by learning-rate scaling alone.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework

    Sep 15, 2026Yunxiang Fu, Meng Lou, Zicheng Liao +1Replay-Based Continual LearningClass-Incremental Learning

  2. To Retain or to Adapt? Generalizing Continual Learning

    Jul 6, 2026Giulia Lanzillotta, Mandana Samiei, Doina Precup +2Replay-Based Continual LearningContinual Learning

  3. Latent-LoRA: Compact Latent-Space Adapters with Gradient-Free Routing for Continual Learning

    Jul 26, 2026Reza Rahimi Azghan, Gautham Krishna Gudur, Giulia Pedrielli +2Replay-Based Continual LearningCatastrophic Forgetting