Features have life history. And we should care
Organizations: University of Tübingen · University of Belgrade · Eindhoven University of Technology
Abstract
Features in language models have life history: they emerge, persist, and die during training, yet the importance of that history remains largely unexplored. We find evidence of a persistent representational backbone, which we identify in Pythia-160M and -410M as the carrier scaffold: sparse features with stable life histories, around which the model's representational structure organises. It has four properties. \emph{(i)}\emph{It assembles early:} features emerge, die, and reorganise faster in the first of training than afterwards, and the scaffold is already largely fixed by then. \emph{(ii)}\emph{It is load-bearing:} joint cross-layer ablation identifies the carriers as far more load-bearing than any count-matched non-scaffold population, a gap invisible to per-firing single-feature methods. \emph{(iii)}\emph{Function precedes direction:} which features will become carriers is already predictable from training-onset firing patterns alone, correctly distinguishing future carriers from non-carriers in of cases, before the geometry has settled. \emph{(iv)}\emph{It seeds subsequent development:} by the end of training, scaffold carriers have recruited of all active features into the scaffold hierarchy. Life history is consistent with a two-phase account of training: selection appears to largely determine the scaffold in the first ; the remaining appears to calibrate geometry around a substrate already set.