Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks
Organizations: VIDRAFT AI Research · QuantumOS, Seoul, Republic of Korea
Abstract
Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model (2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16 larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63% and removing the SSM-family mechanism produces a 3.20% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.
Figures & tables
| Type | Strength | Cost |
| full | Exact all-pairs interactions; no information loss | Quadratic in length |
| sliding | Locality very cheaply, linear in length | Blind beyond the window |
| stronger uniformity differential | Subtracts two attention maps, cancelling common-mode noise | Halves the head dimension |
| linear | Linear-recurrent (Mamba-style) mixing ; linear in length | Approximate—no exact pairwise scores |
| nsa | Gates compressed, selected, and sliding branches for cheap long reach | Structurally complex |
| Item | Value |
| Total parameters | 6.59B |
| Active parameters | 2.98B per token |
| Layers | 49 ( Latin square) |
| Experts | 25, top-7 routing, 1 shared |
| Expert intermediate | 640 |
| Hidden / intermediate | 2048 / 6144 |
| Type | 2K (ms / GB) | 8K (ms / GB) | 32K (ms / GB) |
| full | 0.4 / 0.0 | 1.5 / 0.2 | 13.6 / 0.7 |
| differential | 0.5 / 0.1 | 3.7 / 0.3 | 46.5 / 1.1 |
| sliding | 0.6 / 0.1 | 1.9 / 0.2 | 7.6 / 0.8 |
| nsa | 1.0 / 0.1 | 3.6 / 0.3 | 26.8 / 3.5 |
| hybrid | 1.5 / 0.1 | 7.7 / 0.3 | 74.7 / 3.5 |
| Component | Repository | Config | Tokens | Weight | License |
| English web (edu-filtered) | HuggingFaceFW/ fineweb-edu | sample-100BT | 15.000B | 2.0 | ODC-By |
| Synthetic textbook | HuggingFaceTB/ smollm-corpus | cosmopedia-v2 | 8.000B | 2.0 | ODC-By |
| Math (filtered) | HuggingFaceTB/ finemath | finemath-3plus | 6.002B | 3.5 | ODC-By |
| Code | OpenCoder-LLM/ opc-fineweb- code-corpus | default | 5.001B | 2.5 | MIT |
| Math (web) | open-web-math/ open-web-math | default | 4.004B | 3.5 | source repo |
| Korean web | HAERAE-HUB/ KOREAN-WEBTEXT | default | 2.492B | 2.0 | source repo |
| Item | Value |
| Hardware | NVIDIA B200 16 (2-node FSDP) |
| Total window | 2026-05-30 2026-07-16 ( 46 days ) |
| Final stage | 30 days 11 hours — 11,700 B200-hours |
| Throughput | 32,000 tok/s |
| Steps / tokens | 162,000 / 144.2B |
| Optimizer | AdamW, , |
| Item | Released |
| Weights (annealed base) | [OK] |
| Full architecture source | [OK] |
| Training-data recipe (sources, configs, token counts, weights, tokenizer, EOS) | [OK] |
| Tokenization script | [OK] |
| Training code and launch scripts | [OK] |
| All hyperparameters | [OK] |
| Arm | Layer sequence | Property |
| latin | SGMF GMFS MFSG FSGM | balanced and column-uniform (Latin square) |
| periodic | SGMF SGMF SGMF SGMF | balanced per window, fixed columns (plain cycle) |
| block | SSSS GGGG MMMM FFFF | balanced counts but each type confined to a band |
| Arm | Mean CE | SD | vs. latin | pooled_SD | Verdict |
| latin | 5.28639 | 0.00867 | — | — | Reference |
| periodic | 5.29484 | 0.01611 | +0.16% ( ) | 0.02588 | Null (within noise) |
| block | 5.31754 | 0.01520 | +0.59% ( ) | 0.02475 | Real effect ( pooled SD) |
| homo_F | 5.37540 | 0.01917 | +1.68% ( ) | 0.02976 | Real effect |
| Arm (removed) | Mean CE | SD | vs. latin | pooled SD | Verdict |
| no_S (sliding) | 5.28495 | 0.02197 | ( ) | 0.03340 | Null (within noise) |
| no_G (differential) | 5.28450 | 0.01256 | ( ) | 0.02158 | Null (within noise) |
| no_F (full) | 5.29615 | 0.00284 | ( ) | 0.01291 | Null (within noise) |
| no_M (Mamba-2 / SSM) | 5.39940 | 0.00883 | ( ) | 0.01750 | Real effect |
| baseline latin | 5.28639 | 0.00867 | — | — | Reference (8 seeds) |
| Arrangement | Mechanism distribution across depth | vs. latin |
| latin , periodic | Present at all depths (fully distributed) | 0 (tied) |
| block | Confined to a four-layer band | +0.59% |
| homo_F | Three of four mechanisms absent ; one repeated everywhere | +1.68% |
| Arm | CE (700.9M) | vs. latin | CE (1.514B) | vs. latin |
| latin | 5.28639 | — | 5.26963 | — |
| homo_F (homogeneous) | 5.37540 | +1.68% | 5.40807 | +2.63% |
| no_M (drop SSM) | 5.39940 | +2.14% | 5.43837 | +3.20% |
| Task | acc | acc_norm |
| SciQ | 73.7 | 63.0 |
| PIQA | 66.3 | 65.8 |
| BoolQ | 54.3 | — |
| ARC-Easy | 52.5 | 48.7 |
| WinoGrande | 51.8 | — |
| HellaSwag | 37.2 | 41.0 |
| Task | Score |
| HellaSwag (acc_norm) | 44.6 2.2 |
| COPA | 57.2 2.0 |
| SentiNeg | 55.7 2.5 |
| WiC | 48.8 2.0 |
| BoolQ | 47.8 2.0 |
| Input | Perplexity |
| Natural Korean sentence | 5.4 |
| Same sentence, word order scrambled | 41.3 |
| Random tokens | 2233.4 |
| Item | Location |
| Base model | FINAL-Bench/Aether-7B-5Attn |
| Instruction-tuned | FINAL-Bench/Aether-7B-5Attn-it |
| Second checkpoint, same architecture | FINAL-Bench/AETHER-7B-7Attn-base |
| 11-mechanism extension | FINAL-Bench/Aether-6B-11Attn-base |
| Ablation code | pilot_model.py , pilot_train.py |
| Causal audit | companion paper, Appendix A |