cs.LGSep 28, 2026

Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models

Authors: John Sweeney

Organizations: Sideplane AI

Abstract

Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size ηη on each of sources AA and BB, the weight difference θAB−θBAθ_{AB}-θ_{BA} is, to leading order, η2bABη^2 b_{AB}, where bAB=HBgA−HAgBb_{AB}=H_Bg_A-H_Ag_B is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured θAB−θBAθ_{AB}-θ_{BA}, or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto bABb_{AB} identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on bABb_{AB} decays with further training.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models

    Sep 28, 2026Florian Eichin, Philipp Mondorf, Andrei Mircea +3MemorizationRecurrent Model

  2. How Do Language Models Choose Between Context and Memory?

    Sep 1, 2026Benjamin Shih, John Winnicki, Arianna Cao