Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters
Organizations: Independent Researcher
Abstract
Ternary (1.58-bit) weights are attractive for microcontroller-class language models, but the sub-1M-parameter regime rests mainly on isolated, single-seed comparisons. One prominent example reports that a routed ternary block (convolution, diagonal SSM and sparse attention mixed by a per-token router) beats a parameter-matched full-precision transformer by 22% at 60K parameters, attributing this to inductive bias. We re-run it under one fixed recipe, three seeds per cell, 98 byte-level runs on one laptop. (i) Baseline shape dominates: at a 16M-byte budget, param-matched transformers span 22.6% in validation loss purely by depth/width choice - far more than any architecture effect we measure there - and the best-shaped transformer ties the routed model, so the published margin is at least partly a baseline-shape effect; the ordering of shapes reverses with budget, so no single fixed shape can be trusted. (ii) At 130M bytes the routed model does win, by 22.2-24.0% over the three transformer shapes we evaluate there - but a plain gated diagonal-SSM block beats it by a further 9.1%, and the routed model's own router puts most of its weight on its recurrent pathway, so the gain does not require routing. (iii) The ternary penalty differs by architecture at the larger budget (+5.3% best transformer vs. +19.5% routed, +28.1% gated SSM), but we cannot attribute that to architecture alone: our transformers keep learned positional embeddings in full precision, 11-22% of their parameters, so they are less quantized than the models they are compared with. (iv) A 90/10 full-precision-then-ternary schedule beats all-ternary training, but only at a stage-2 learning rate about 10x the pretraining peak; at a conventional fine-tuning rate it looks 15.3% worse, reversing the conclusion. The from-scratch baseline was not itself learning-rate tuned, which bounds (iii) and (iv). Code and run logs released.
Figures & tables
| 60K comparison | 944K comparison | |
|---|---|---|
| Result (routed ternary vs. FP32 transformer) | 6.31 vs. 8.12 ppl (win) | 1.0545 vs. 0.9337 nats/byte (loss) |
| Training tokens | 3M (3,000 steps 16 64) | 2B (30,000 256 256) |
| Tokens per parameter | 50 | 2,100 |
| Corpus | 0.5 MB TinyStories slice | 1.7 GB TinyStories (reported) |
| Sequence length | 64 | 256 |
| LR schedule | constant | warmup 1000, cosine |
| Experiment | cells | seeds | runs |
|---|---|---|---|
| Main grid (routed, transformer) | 8 | 3 | 24 |
| Shape sweep (5 shapes @16M, 3 @130M) | 12 | 3 | 36 |
| Gated SSM | 4 | 3 | 12 |
| FP-init arms (2 architectures 2 stage-2 LRs, incl. stage 1) | 6 | 3 | 18 |
| Stage-2 LR sweep | 10 | 1 | 10 |
| Total | 98 distinct |
| System | configuration | params | FP kept | what stays full precision |
|---|---|---|---|---|
| Routed 3-pathway | , 4 layers, , top- | 60,800 | 2.3% | SSM scalars, LayerNorms |
| Gated SSM (ours) | , 3 layers, | 64,196 | 1.6% | SSM scalars, LayerNorms |
| Transformer | , 1 layer, , 4 heads | 61,880 | 22.0% | + positional embedding |
| Transformer | , 2 layers, , 4 heads | 59,112 | 16.2% | + positional embedding |
| Transformer | , 3 layers, , 4 heads | 61,888 | 14.0% | + positional embedding |
| Transformer | , 4 layers, , 4 heads | 59,640 | 12.9% | + positional embedding |
| 16M bytes | 130M bytes | ||||
|---|---|---|---|---|---|
| System | FP32 | ternary | FP32 | ternary | penalty @130M |
| Gated SSM (ours) | 2.704 .041 | 3.376 .093 | 1.516 .012 | 1.942 .018 | +28.1% |
| Routed 3-pathway (Atome) | 2.808 .062 | 3.290 .065 | 1.668 .024 | 1.993 .032 | +19.5% |
| Transformer | 2.835 .016 | 3.426 .034 | 2.195 .003 | 2.311 .003 | +5.3% |
| Transformer | 3.080 .031 | 3.752 .048 | 2.169 .016 | 2.320 .002 | +7.0% |
| Transformer | 3.239 .152 | 3.870 .094 | 2.143 .028 | 2.328 .004 | +8.6% |
| stage-2 LR | 3e-5 | 1e-4 | 3e-4 | 1e-3 | 3e-3 | arm (3 seeds) | loss | vs. scratch | |
|---|---|---|---|---|---|---|---|---|---|
| routed | 2.758 | 2.335 | 2.110 | 1.944 | 1.911 | routed, 1e-4 | 5.2% | ||
| transformer | 2.701 | 2.474 | 2.348 | 2.298 | 2.265 | routed, 3e-3 | 3.4% | ||
| transf., 1e-4 | 1.2% | ||||||||
| transf., 3e-3 | 0.4% |
| Claim | Evidence | Status |
|---|---|---|
| Baseline depth/width choice can decide an architecture verdict at 60K | 5-shape sweep, 2 budgets | supported |
| The published 60K advantage is explained by training budget (H-budget) | 16M vs. 130M bytes | not supported |
| Routing is required for the routed block’s advantage | gated SSM comparison | not supported |
| A simple gated recurrence suffices to match or beat it | gated SSM, 2 budgets | supported |
| Recurrence per se is the causal ingredient | not isolated (FFN, mixing differ) | untested |
| The ternary penalty differs by architecture at 130M | 3 families, 2 budgets | supported |