Forking: Sudden Overfitting Under Replay
Organizations: MetaCircle (元环智能) · Tsinghua University · Peking University · Shanghai Qizhi Institute
Abstract
This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoGPT setting and reproduce it in a DeepSeek-style model with Engram. Mechanistically, repeated updates sharpen the continuations observed in training while suppressing the probability of unseen continuations, whose loss grows with each pass. The n-gram module creates weakly interacting context-specific subspaces, amplifying this effect. Low-frequency contexts contribute most of the gap, whereas larger training budgets and heavily crowded tables suppress it. We also observe forking in short-budget, heavily repeated SFT and RL-like regimes. The contributions of this paper are twofold: (1) Forking reveals yet another curious phenomenon in deep learning, in addition to grokking and double descent. (2) Forking is an unexpected and unpleasant by-product of tricks proposed by autoresearch agents. While these agents produce an enormous number of results that seem useful, we should always be careful with their results.
Figures & tables
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| \toprule Component | Configuration |
| \midrule Backbone | 8 layers; hidden dimension 768; 6 attention heads; head dimension 128. |
| Input and batch | Sequence length 2048; 72 sequences or 147456 tokens per update; gradient accumulation 1. |
| Tokenizer | Shared tokenizer; vocabulary 8192; built from the first ten shards. |
| N-gram tables | Two tables per order and injection layer; each has 524288 rows of dimension 384. Bigram at layers 2, 4, 6, 8; trigram at layers 2, 6, 8. |
| Optimizers | Muon for matrix parameters, base learning rate 0.04; grouped AdamW for token/value embeddings, output and scalar parameters; RMSProp for bigram/trigram tables, learning rate 0.6, second-moment decay 0.999, no weight decay. |
| Schedule | Archived baseline schedule: Muon and ordinary AdamW warmdown begins at 5% and 35% of the step budget, respectively; final multiplier 0.05. N-gram table learning rates remain constant. |
| \toprule Component | Configuration |
| \midrule Attention backbone | 8 layers; hidden width 768; 6 heads of dimension 128; full causal attention in every layer. |
| Feed-forward blocks | Two linear layers with a expansion to width 3072 and a GELU activation. |
| Normalization and positions | Pre-LayerNorm in each block and a final LayerNorm, ; learned absolute position embeddings. |
| Embeddings and regularization | Vocabulary 8192; tied token/output weights; dropout 0. Attention and MLP linear layers and LayerNorm include biases; the output head does not. |
| N-gram memory | One hash table per order, bigram and trigram; full-width lookup vectors added at the input. No extra unigram or fourgram memory, gates, or multi-hash concatenation. |
| Sequence and batch | 2048 prediction positions per sequence; 72 sequences per update; 147,456 prediction tokens; one device batch per update. |
| \toprule Condition | Intervention | Reported observation |
| \midrule No-gram | Remove the n-gram module | Raw gap 0.233 at step 2000; ordinary train–validation differences remain. |
| N-gram | Input injection; table LR | Raw gap 5.676 at step 2000, with replay-related steps. |
| Low-frequency mask | Mask low-count reads in training and evaluation | removes about 72% of net gap at step 1000; higher thresholds further reduce it (raw gap 0.101 at ). |
| High-frequency mask | Mask high-count reads | leaves the step-1000 raw gap at 2.824 (control 2.724), a much smaller change than low-frequency masking. |
| Hash reseed | Change context-to-row mapping at each replay boundary; retain parameters and optimizer | Raw gap 0.069 at step 1000 when reseeded at steps 338 and 675 (1.354 with one reseed); substantial suppression of replay accumulation. |
| Freeze table | Stop table updates at step 675 | 94% of the control gap remains at step 2022 (5.386 versus 5.733) and 72% after ten passes; the gap keeps growing, more slowly than in control. Frozen at step 338, the step-1000 gap is 3.452. |
| \toprule context | class | share of positions | share of excess loss |
|---|---|---|---|
| \midrule bigram | novel continuation | 31.26% | 72.45% |
| bigram | seen continuation | 68.74% | 27.55% |
| trigram | novel continuation | 65.67% | 102.02% |
| trigram | seen continuation | 34.33% | % |
| \bottomrule |
| \toprule Context | Gap exponent | Novelty exponent | ||
|---|---|---|---|---|
| \midrule Bigram | 0.251 | 0.997 | 0.229 | 0.981 |
| Trigram | 0.466 | 0.843 | 0.212 | 0.985 |
| \bottomrule |
| \toprule Context | Step | rms | Relative | Holdout | ||
| \midrule Bigram | 670 | 4.87 | 0.137 | 8.2% | 0.169 | |
| 1000 | 9.58 | 0.233 | 6.3% | 0.269 | ||
| 1340 | 10.86 | 0.16 | 0.205 | 4.1% | 0.228 | |
| 2000 | 10.90 | 1.89 | 0.202 | 3.0% | 0.233 | |
| Trigram | 670 | 2.35 | 0.063 | 10.7% | 0.060 | |
| 1000 | 5.81 | 0.155 | 9.9% | 0.132 |
| \toprule Component | Configuration |
| \midrule Decoder and embeddings | 12 layers; hidden width 768; vocabulary 8192; sequence length 2048; separate token and output embeddings; dropout 0. |
| MLA-lite attention | 12 heads; query latent rank 256 and key/value latent rank 128. Per head: 64 non-rotary query/key dimensions, 32 rotary dimensions, and 64 value dimensions; causal attention; RoPE base 10,000. |
| Feed-forward layers | Layer 1: dense SwiGLU with intermediate width 2048. Layers 2–12: 16 routed experts, top-2 selection, and two shared experts; each expert has intermediate width 512. |
| Expert balancing | Sigmoid router scores, normalized over the selected experts; auxiliary-loss-free correction bias with update rate . |
| Residual connections | Four mHC streams with learned pre-, post-, and residual mixing; 20 Sinkhorn iterations. RMSNorm is used in latent attention, memory fusion, and the output pathway. |
| Memory locations | Engram at layers 2 and 6; separate parameters at each layer; bigram and trigram retrieval. |
| \toprule Component | Configuration |
| \midrule Non-table parameters | AdamW; learning rate ; ; . Weight decay 0.1 on matrix and convolution weights, excluding token/output embeddings; biases and one-dimensional parameters have no decay. |
| Lookup tables | Bias-corrected, momentum-free RMSProp; learning rate ( the backbone rate); second-moment decay 0.99; ; no weight decay. |
| Learning-rate schedule | Both rates increase linearly from to over updates 1–100, then remain constant; no warmdown or later decay. |
| Numerics and execution | BF16 forward computation; FP32 model parameters, optimizer states, and cross-entropy logits; activation checkpointing; eager execution; one NVIDIA H200 per run, with the official mHC kernels. |
| Evaluation | Current-batch training loss and fixed validation loss every 10 updates; 200 recorded points per run; no smoothing. |
| \bottomrule |
| \toprule Source | Table | Train | Validation | Gap |
|---|---|---|---|---|
| \midrule ID | Trainable | 1.913 | 4.376 | 2.463 |
| ID | Frozen | 3.030 | 3.372 | 0.341 |
| Magicoder | Trainable | 0.619 | 1.543 | 0.924 |
| Magicoder | Frozen | 1.103 | 1.301 | 0.198 |
| OpenR1 | Trainable | 0.688 | 1.366 | 0.678 |
| OpenR1 | Frozen | 1.039 | 1.197 | 0.158 |
| \toprule Source | Table | Train | Validation | Gap |
|---|---|---|---|---|
| \midrule RLVR-MATH | Trainable | 0.009 | 0.518 | 0.509 |
| RLVR-MATH | Frozen | 0.014 | 0.447 | 0.434 |
| \bottomrule |