cs.LGJul 16, 2026

Gate-Zero Growth: A Geometric Framework for Function-Preserving Continual Learning

Authors: Dante Lok

Organizations: Votee AI · Beever AI

Abstract

We introduce \emph{gate-zero growth}, a function-preserving (FP) operator for continual learning that adds new residual blocks through a zero-initialised gate. Under a transversality condition, gate-zero growth induces \emph{rank separation} in the functional Jacobian: old directions are unchanged, new-weight directions are exactly flat at the growth point, and new gate directions are the only first-order source of new functional variation. As gates open during continual learning, function drift is O(α2)O(\|\boldsymbolα\|^2) and Jacobian leakage O(α)O(\|\boldsymbolα\|_\infty), giving a controlled departure from the FP locus. On a 300M857M300\mathrm{M}\to857\mathrm{M} Transformer adapted from WikiText-103 to BookCorpus, gate-zero growth reaches near-zero old-domain forgetting (ΔA<0.1Δ_A < 0.1) under both exact-preservation (Isolation) and joint-frontier (Freeze-Nothing) operating points, while a non-FP control (GstackG_{\text{stack}}) suffers an order-of-magnitude larger forgetting under the same recipe. The same geometric analysis covers LoRA, ReZero, and zero-init adapter constructions, establishing gate-zero growth as the canonical instance of a shared local geometry that governs safe capacity activation in CL.

Explore similar work

CardsList