Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models
Organizations: Sideplane AI
Abstract
Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size on each of sources and , the weight difference is, to leading order, , where is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured , or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on decays with further training.
Figures & tables
| Object | Role | Origin |
|---|---|---|
| effect of swapping and : | inherited | |
| one scalar predicting which order has lower held-out loss | inherited | |
| (commutator memory) | token-level decomposition of ; | new |
| order-gap closure | fraction of the AB/BA NLL gap removed by an intervention | new |
| which order produced each of two endpoints | new | |
| bracket-only step; beats both orders under Prop. 1 | partly inherited |
| Model | Precision | 80% mass mean/med. | Endpoint/resampled overlap | Random overlap |
|---|---|---|---|---|
| Llama-3.2-1B | fp32 | 1.47% / 1.47% | 99% / 97% | 35% |
| Qwen-2.5-1.5B | fp32 | 0.99% / 1.16% | 99% / 97% | 39–40% |
| Qwen-3-4B | fp32 | 0.30% / 0.042% | 82% / 93% | 36–49% |
| Model | Condition | Winsor. mean | Median | 95% CI | Gap shrinks |
|---|---|---|---|---|---|
| Qwen-3-4B | Harmful | +31.2% | +32.0% | [+21.4%, +39.7%] | 75.6% |
| Qwen-3-4B | Random | +1.2% | +0.2% | [–5.8%, +8.7%] | 50.0% |
| Qwen-2.5 † | Harmful | +48.6% | +47.9% | [+40.5%, +57.1%] | 96.4% |
| Qwen-2.5 † | Random | –0.006% | –0.003% | [–0.03%, +0.02%] | 39.3% |
| Model (precision) | correct | Wilson 95% CI |
|---|---|---|
| Llama-3.2-1B (1B, fp32) | 14/18 = 77.8% | |
| Qwen-2.5-1.5B (1.5B, fp32) | 16/18 = 88.9% | |
| Qwen-3-4B (4B, bf16) | 18/18 = 100.0% | |
| Llama-3.1-8B (8B, bf16) | 18/18 = 100.0% | |
| Combined | 66/72 = 91.7% |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Pair | mean | flat CI | cluster CI | median | wrong-sign mean | random mean |
| fq:flan | 0.675 | [0.561, 0.767] | [0.425, 0.772] | 0.796 | ||
| fq:share | 0.454 | [0.157, 0.703] | [0.126, 0.536] | 0.739 | ||
| fq:ultra | 0.647 | [0.477, 0.798] | 0.778 | |||
| flan:share | 0.545 | [0.247, 0.776] | 0.807 | |||
| flan:ultra | 0.720 | [0.645, 0.788] | [0.455, 0.754] | 0.785 | ||
| ultra:share | 0.454 | [0.219, 0.665] | 0.735 |
| Pair | paired difference (mean) | CI 95% | flat diagnostic | direction |
|---|---|---|---|---|
| fq:flan | positive | mitigation | ||
| fq:share | negative | excludes 0 | positive | mitigation |
| fq:ultra | negative | excludes 0 | positive | mitigation |
| flan:share | — | includes 0 | inconclusive | no effect |
| flan:ultra | negative | excludes 0 | positive | mitigation |
| ultra:share | negative | excludes 0 | positive | mitigation |
| Pair | rel. err. | ||
|---|---|---|---|
| false_qa_vs_flan_v2_p3 | |||
| false_qa_vs_sharegpt | |||
| false_qa_vs_ultrachat | |||
| flan_v2_p3_vs_sharegpt | |||
| flan_v2_p3_vs_ultrachat | |||
| ultrachat_vs_sharegpt |
| Pair | Single-step closure (%) | Spearman( , ) | |
|---|---|---|---|
| math-vs-legal | 0.042 | +80.3% | 0.833 |
| code-vs-legal | 0.013 | +28.4% | 0.800 |
| news-vs-legal | 0.037 | +23.9% | 0.830 |
| code-vs-news | 0.004 | +5.9% | 0.832 |
| news-vs-biomedical | 0.001 | 0.844 | |
| code-vs-biomedical | 0.002 | 0.821 |
| Model | Dtypes | Top-1% | Top-5% | Sign | |
|---|---|---|---|---|---|
| Qwen-3-4B | bf16/fp32 FD | 53.0 (54.5 med.)% | 51.0 (54.3 med.)% | 52.4% | |
| Qwen-2.5-1.5B | fp32/fp32 FD | 57.3 (57.1 med.)% | 46.8 (47.5 med.)% | 70.6 1.9% |
| Model | Scale | correct | Wilson 95% CI |
|---|---|---|---|
| Llama-3.2-1B (bf16 storage) | 1B | 14/18 = 77.8% | |
| Llama-3.2-1B (fp32) | 1B | 14/18 = 77.8% | |
| Qwen-2.5-1.5B (fp32) | 1.5B | 16/18 = 88.9% | |
| Qwen-2.5-1.5B (bf16 storage) | 1.5B | 16/18 = 88.9% | |
| Qwen-3-4B (bf16 storage) | 4B | 18/18 = 100.0% | |
| Llama-3.1-8B (bf16 storage) | 8B | 18/18 = 100.0% |
| Model | Pair | Gini | pct for | pct for |
|---|---|---|---|---|
| Qwen-2.5-1.5B (fp32) | false_qa vs. flan_v2_p3 | |||
| Qwen-2.5-1.5B (fp32) | ultrachat vs. sharegpt | |||
| Qwen-3-4B (bf16) | false_qa vs. flan_v2_p3 | |||
| Qwen-3-4B (bf16) | ultrachat vs. sharegpt | |||
| Qwen-3-4B (bf16) | false_qa vs. ultrachat | |||
| Qwen-3-8B (bf16) | false_qa vs. flan_v2_p3 |
| Pair | CI excludes 0 | paired difference CI 95% | direction |
|---|---|---|---|
| evol_instruct_vs_truthful_qa | yes | excludes 0 | mitigation |
| evol_instruct_vs_false_qa | no | includes 0 | no effect |
| truthful_qa_vs_flan_v2_p3 | yes | excludes 0 | mitigation |
| Pair | bracket mean | median | random | wrong-sign | first-order |
|---|---|---|---|---|---|
| evol:fq | |||||
| evol:flan | |||||
| evol:share | |||||
| truth:flan | |||||
| truth:share | |||||
| truth:ultra |
| Pair | flat-bootstrap CI 95% | seed-cluster CI 95% |
|---|---|---|
| false_qa_vs_flan_v2_p3 | [0.561, 0.767] | [0.425, 0.772] |
| false_qa_vs_sharegpt | [0.157, 0.703] | [0.126, 0.536] |
| false_qa_vs_ultrachat | [0.477, 0.798] | |
| flan_v2_p3_vs_sharegpt | [0.247, 0.776] | |
| flan_v2_p3_vs_ultrachat | [0.645, 0.788] | [0.455, 0.754] |
| ultrachat_vs_sharegpt | [0.219, 0.665] |
| Triplet | Seed mean | Seed mean | Cells | Triplet mean |
|---|---|---|---|---|
| code, legal, news | ||||
| code, biomedical, legal | ||||
| news, biomedical, math | ||||
| Grand mean ( cells) | ||||
| Pair | top-K | non-top | null | CG topK | CG full | ||||
|---|---|---|---|---|---|---|---|---|---|
| code/news | |||||||||
| code/math | |||||||||
| legal/bio | |||||||||
| mean |
| Model | Regime | |
|---|---|---|
| Qwen-3-4B | cube-root | |
| Qwen-2.5-1.5B | cube-root | |
| Llama-3.1-8B | cube-root | |
| Llama-3.2-1B | cube-root | |
| GPT-2 (reference) | strongly powered |
| Model | Precision | Method | Harmful (med.) | Random (med.) | Sep. |
|---|---|---|---|---|---|
| Qwen-3-4B | bf16 | Loss reweighting | +32.0% | +0.2% | ✓ |
| Qwen-3-4B | bf16 | LR scaling † | +31.2% | +1.7% | ✓ |
| Qwen-3-4B | bf16 | Projection † | +14.6% | –5.1% | |
| Qwen-3-4B | bf16 | Regularizer † | +8.3% | +2.9% | |
| Qwen-2.5-1.5B | fp32 | LR scaling | +47.9% | –0.003% | ✓ |
| Metric | Value |
|---|---|
| Plotted points | |
| of the fit (mean over seeds) | 0.994 |
| Pooled predicted-vs-actual | 0.990 |
| Best-fit slope in pooled plot | 0.706 |
| Mean magnitude ratio at |
| Slice (run) | Tokens, and the readout of the colored ones |
|---|---|
| code/news | Ask H N : Should I include my GPA and /or transcripts when applying for jobs ? Ask : (rank ) |
| code/biomedical | Comput ers are a ubiquitous part of the amb ulatory health care environment . ulatory : (rank ); of : (rank ) |
| code/news (iterative run) | Facebook Dating launch blocked in Europe after it fails to show privacy workings Facebook : (rank ); its NLL is under AB and under BA, and after bracket correction ( of its NLL gap closed) |
| Method | Qwen-3-4B | Qwen-2.5-1.5B | Llama-3.1-8B |
|---|---|---|---|
| Bracket | 72% (13/18) | 100% (18/18) | 100% (18/18) |
| Grad Cosine | 11% (2/18) | 33% (6/18) | 22% (4/18) |
| Grad Norm | 11% (2/18) | 56% (10/18) | 22% (4/18) |
| Random | 44% (8/18) | 50% (9/18) | 39% (7/18) |
| Model | Storage/HVP | Gini | recall@1%V | gap/gap 1 | ||
|---|---|---|---|---|---|---|
| Qwen-3-4B | bf16/fp32 | 1 | 0.9998 | 100% | 1.0 | 1 |
| 2 | 0.9998 | 100% | 1.5 | 4 | ||
| 4 | 0.9998 | 100% | 2.0 | 16 | ||
| 8 | 0.9997 | 100% | 3.1 | 64 | ||
| Qwen-2.5-1.5B | fp32 | 1 | 0.9982 | 99% | 1.0 | 1 |
| 2 | 0.9982 | 99% | 4.0 | 4 |
| Pair | Seed | |||||
|---|---|---|---|---|---|---|
| code/news | legal | 8042 | 0.663 | 0.286 | 0.175 | 0.134 |
| code/news | legal | 9042 | 0.530 | 0.291 | 0.219 | 0.113 |
| code/legal | news | 8042 | 2.513 | 0.433 ∘ | 0.309 ∘ | 0.170 ∘ |
| code/legal | news | 9042 | 2.337 | 0.348 ∘ | 0.257 ∘ | 0.157 ∘ |
| code/biomedical | news | 8042 | 0.721 | 0.583 | 0.446 | 0.292 |
| code/biomedical | news | 9042 | 0.577 | 0.576 | 0.439 | 0.283 |
| Experiment family | Hardware | GPU-hours | Notes |
|---|---|---|---|
| Sparsity (4 models) | RTX 4090 / Pro 6000 | 2 | R8b, R9, R13_sparsity, R14_sparsity, R15_sparsity, R16 |
| Baselines (4 models) | RTX 4090 / A40 | 1 | R8b, R9, R13 |
| SFT interventions | RTX Pro 6000 / RTX 4090 | 3 | R14a, R14b, R15 (full fp32 for Qwen-2.5) |
| Endpoint correction / multi-step | A40 / L40 / Pro 6000 | 5 | R26, R27, R28, R19, R20 |
| Iterative bracket correction | L40 | 3 | R29, R30 (instrumented per-position NLL) |
| DPO sparsity (3 models) | A40 / Pro 6000 | 2 | R50, R52, R60 |
| Condition | Parameter closure | Log-prob closure | Loss closure |
|---|---|---|---|
| Bracket edit | |||
| Random norm-matched | |||
| Wrong sign | |||
| First-order norm-matched | |||
| Unmatched rollouts (negative control) |
| identity cosine | relative residual | |||
|---|---|---|---|---|
| 2 | ||||
| 4 | ||||
| 8 | ||||
| 16 |