cs.MAOct 7, 2026

Decentralized collaborative continual learning: A multi-objective minimization-based technique

Authors: Yara Zgheib, Marc Antonini, Roula Nassif

Organizations: Universit´e Cˆote d’Azur, I3S Laboratory, CNRS, France

Abstract

In this work, we formulate decentralized continual learning within a multi-objective optimization framework. For a given inference task t (corresponding to a common minimizer shared by the cost functions of all agents), agents collecting data in a distributed and streamed manner are only allowed to perform local computations and to exchange information with neighboring agents over the underlying communication graph. As tasks evolve sequentially over time, agents must adapt to newly arriving tasks while retaining knowledge acquired from previously learned ones. This requirement leads to the wellknown stability plasticity dilemma, where stability refers to the ability to retain previous knowledge, while plasticity refers to the ability to learn and adapt to new tasks. To address the stability challenge, agents store subsets of samples from past tasks in local memory buffers. Then, through an appropriate multiobjective formulation, the stored information is incorporated into the learning process so that parameter updates account jointly for the current task and previously learned tasks. The proposed decentralized continual learning approach is analyzed in the mean square error sense under general assumptions on the individual cost functions and gradient noise processes. The analysis reveals that cooperation among agents improves the performance of continual learning. In particular, by exchanging information with neighboring agents, decentralized collaborative learning can exploit the diversity of locally observed data and memory buffers to improve the network average mean-square deviation (MSD) across tasks. Finally, simulations illustrate the theoretical findings and the effectiveness of the method in reducing forgetting and improving the average MSD across tasks.

Figures & tables

Explore similar work

Apr 20, 2026cs.LG

Task Switching Without Forgetting via Proximal Decoupling

In continual learning, the primary challenge is to learn new information without forgetting old knowledge. A common solution addresses this trade-off through regularization, penalizing changes to parameters critical for previous tasks. In most cases, this regularization term is directly added to the training loss and optimized with standard gradient descent, which blends learning and retention signals into a single update and does not explicitly separate essential parameters from redundant ones. As task sequences grow, this coupling can over-constrain the model, limiting forward transfer and leading to inefficient use of capacity. We propose a different approach that separates task learning from stability enforcement via operator splitting. The learning step focuses on minimizing the current task loss, while a proximal stability step applies a sparse regularizer to prune unnecessary parameters and preserve task-relevant ones. This turns the stability-plasticity into a negotiated update between two complementary operators, rather than a conflicting gradient. We provide theoretical justification for the splitting method on the continual-learning objective, and demonstrate that our proposed solver achieves state-of-the-art results on standard benchmarks, improving both stability and adaptability without the need for replay buffers, Bayesian sampling, or meta-learning components.
Jul 6, 2026stat.ML

To Retain or to Adapt? Generalizing Continual Learning

The Continual Learning (CL) literature has long been driven by the goal of mitigating catastrophic forgetting. This objective rests on a pervasive, often unstated assumption: that a lifelong learner should approximate the Joint-Task Learning (JTL) solution and retain all previously acquired knowledge. We challenge this retention-centered premise, arguing that in non-stationary environments prioritizing retention can impede real-time adaptation. Shifting the focus to the Average Lifelong Error (ALE), we formalize CL as an online optimization problem governed by the interaction between environmental and learning dynamics. We introduce Transfer Efficiency as a quantitative measure of the tension between Instability, the bias inherited from conflicting past experience, and Transient Error, the optimization cost of learning new tasks from scratch. Under mild convergence conditions, holding across linear and neural network models, this decomposition yields a Critical Task Duration: a closed-form threshold beyond which historical knowledge transitions from a warm-start advantage to an optimization liability whenever retention induces a positive stationary bias. We validate these theoretical predictions on continual image classification and reinforcement learning benchmarks. Finally, by connecting continual learning to the online learning framework of predictable sequences, we show that JTL is only one instance of a broader family of objectives, and we propose a new general class of continual learning algorithms, which we call Predictive Continual Learning. Predictive CL algorithms optimize expected future performance under an explicit, dynamically updated model of future tasks. As a proof of concept, we analyze a Window algorithm that interpolates between JTL and Independent-Task Learning (ITL), outperforming both under controlled distributional drift.
Apr 15, 2026cs.LG

From Order to Distribution: An Exact Operator Framework for Forgetting in Continual Learning

A central challenge in continual learning is forgetting: the loss of performance on previously learned tasks after learning new ones. Prior theory has analyzed forgetting under random orderings of fixed task collections in overparameterized linear regression. We shift the focus from task order to task distribution, asking how its structure determines forgetting. In the linear setting with a shared solution, i.i.d. task sampling, and sequential exact fitting, we derive an exact operator identity expressing historical forgetting directly in terms of the task distribution. Building on this identity, we establish an exponential decay guarantee for expected historical forgetting under every fixed task distribution in finite dimensions, characterize its asymptotic behavior, and relate decay to the distribution's coverage of observable directions. For an individual learned task, we show that subsequent tasks can collectively support recovery without exact revisits. We derive a lower bound on recovery time and construct a task distribution attaining its inverse-coverage scaling.