| Experiment | Model | Mini-batch | Horizon and seeds |
|---|
| Figure 2 | One-layer causal transformer with embedding dimension 512 , 8 attention heads, feed-forward width 256 , and dropout 0.2 . The embedding, transformer layer, and output classifier are all trained. The model has approximately 11.49 M trainable parameters. | 1,024×35 | 100 epochs; seeds 0,1,2 for the optimizer comparison and seed 0 for the per-frequency-group runs. |
| Figure 3 (a) | One-layer transformer with embedding dimension 256 , one attention head, feed-forward width 128 , and dropout 0.2 . The token embedding and complete transformer encoder are frozen; only the softmax classification head is trained. The model has approximately 5.42 M parameters, of which 2.55 M are trainable. | 1,024×35 | 100 epochs; seeds 0,1,2 . |
| Figure 3 (b) | Bigram model obtained by removing attention. It consists of a fixed 9,922×256 token-embedding matrix followed by a trainable, bias-free 256×9,922 output classifier. It has approximately 5.08 M parameters, of which 2.54 M are trainable. | 1,024×35 | 100 epochs; seeds 0,1,2 . |
| Figure 3 (c,d) | Collapsed softmax unigram model. Every input token is replaced by the same index, leaving one trainable vector of 9,922 logits. The logits are initialized to zero. | 1,024×32 | 200 epochs; seeds 0,1,2 . |
| Figure 1 | Deterministic softmax unigram model with d=10,000 trainable logits, initialized to zero, and target probabilities pi∝i−1 . | Full batch | 600 gradient steps; seed 0 . |