Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model (≈2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a 7×7 Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a 4×4 Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16× larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63% and removing the SSM-family mechanism produces a 3.20% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.
Figures & tables
Type
Strength
Cost
full
Exact all-pairs interactions; no information loss
Quadratic in length
sliding
Locality very cheaply, linear in length
Blind beyond the window
stronger uniformity differential
Subtracts two attention maps, cancelling common-mode noise
Halves the head dimension
linear
Linear-recurrent (Mamba-style) mixing ; linear in length
Approximate—no exact pairwise scores
nsa
Gates compressed, selected, and sliding branches for cheap long reach
Structurally complex
Table 1: Five base sequence-mixing mechanisms used in Aether. Each mechanism provides a distinct inductive bias, computational profile, and failure mode; the motivation for heterogeneous architectures is that no single mechanism dominates across all dimensions.
Item
Value
Total parameters
6.59B
Active parameters
≈ 2.98B per token
Layers
49 ( 7×7 Latin square)
Experts
25, top-7 routing, 1 shared
Expert intermediate
640
Hidden / intermediate
2048 / 6144
Table 2: Architectural specification of Aether-7B-5Attn. Active parameters denote the average number of parameters participating in each token computation under MoE routing.
Type
2K (ms / GB)
8K (ms / GB)
32K (ms / GB)
full
0.4 / 0.0
1.5 / 0.2
13.6 / 0.7
differential
0.5 / 0.1
3.7 / 0.3
46.5 / 1.1
sliding
0.6 / 0.1
1.9 / 0.2
7.6 / 0.8
nsa
1.0 / 0.1
3.6 / 0.3
26.8 / 3.5
hybrid
1.5 / 0.1
7.7 / 0.3
74.7 / 3.5
Table 3: Per-mechanism computational cost profile. Values report prefill latency (ms) and peak memory usage (GB) for an isolated layer at different context lengths.
Component
Repository
Config
Tokens
Weight
License
English web (edu-filtered)
HuggingFaceFW/ fineweb-edu
sample-100BT
15.000B
2.0
ODC-By
Synthetic textbook
HuggingFaceTB/ smollm-corpus
cosmopedia-v2
8.000B
2.0
ODC-By
Math (filtered)
HuggingFaceTB/ finemath
finemath-3plus
6.002B
3.5
ODC-By
Code
OpenCoder-LLM/ opc-fineweb- code-corpus
default
5.001B
2.5
MIT
Math (web)
open-web-math/ open-web-math
default
4.004B
3.5
source repo
Korean web
HAERAE-HUB/ KOREAN-WEBTEXT
default
2.492B
2.0
source repo
Table 4: Training corpus composition. We report the source repository, configuration, token count, sampling weight, and license information for each component used in the data mixture.
Item
Value
Hardware
NVIDIA B200 × 16 (2-node FSDP)
Total window
2026-05-30 → 2026-07-16 ( ≈ 46 days )
Final stage
30 days 11 hours — ≈ 11,700 B200-hours
Throughput
≈ 32,000 tok/s
Steps / tokens
162,000 / 144.2B
Optimizer
AdamW, β=(0.9,0.95) , ε=10−8
Table 5: Training compute, optimization configuration, and resource utilization for Aether-7B-5Attn.
Table 6: Artifacts released with Aether-7B-5Attn. The release includes model checkpoints, source code, training recipes, and supporting resources intended to facilitate independent verification and reproduction.
Figure 5: Measured per-mechanism prefill cost at 2K/8K/32K.
Arm
Layer sequence
Property
latin
SGMF GMFS MFSG FSGM
balanced and column-uniform (Latin square)
periodic
SGMF SGMF SGMF SGMF
balanced per window, fixed columns (plain cycle)
block
SSSS GGGG MMMM FFFF
balanced counts but each type confined to a band
Table 7: Placement-axis ablation. All arms contain the identical 4+4+4+4 mechanism multiset and identical parameter counts; only the depth-wise arrangement of mechanisms differs. The Latin-square and periodic schedules maintain distributed exposure, whereas the block schedule concentrates each mechanism within a single depth band.
Arm
Mean CE
SD
Δ vs. latin
2⋅ pooled_SD
Verdict
latin
5.28639
0.00867
—
—
Reference
periodic
5.29484
0.01611
+0.16% ( Δ=0.00845 )
0.02588
Null (within noise)
block
5.31754
0.01520
+0.59% ( Δ=0.03115 )
0.02475
Real effect ( 2.5× pooled SD)
homo_F
5.37540
0.01917
+1.68% ( Δ=0.08901 )
0.02976
Real effect
Table 8: Placement-axis results and homogeneous reference (eight seeds per arm; confirmed 2026-07-23). The Latin, periodic, and block arms share identical mechanism composition and parameter counts, differing only in depth-wise arrangement. The homogeneous reference ( homo_F ) represents the limiting case in which a single mechanism occupies every layer.
Arm (removed)
Mean CE
SD
Δ vs. latin
2× pooled SD
Verdict
no_S (sliding)
5.28495
0.02197
−0.00144 ( −0.03% )
0.03340
Null (within noise)
no_G (differential)
5.28450
0.01256
−0.00189 ( −0.04% )
0.02158
Null (within noise)
no_F (full)
5.29615
0.00284
+0.00976 ( +0.18% )
0.01291
Null (within noise)
no_M (Mamba-2 / SSM)
5.39940
0.00883
+0.11301 ( +2.14% )
0.01750
Real effect
baseline latin
5.28639
0.00867
—
—
Reference (8 seeds)
Table 9: Composition-axis leave-one-out ablation (four seeds per arm; confirmed 2026-07-23). Each row removes one mechanism from the heterogeneous stack while re-matching parameters and cycling the remaining three mechanisms throughout depth. Results are reported relative to the latin baseline.
Arrangement
Mechanism distribution across depth
vs. latin
latin , periodic
Present at all depths (fully distributed)
0 (tied)
block
Confined to a four-layer band
+0.59%
homo_F
Three of four mechanisms absent ; one repeated everywhere
+1.68%
Table 10: Monotonic relationship between depth-wise mechanism confinement and validation-loss degradation. Arrangements are ordered by the extent to which individual mechanisms are restricted to particular depth ranges. Performance deteriorates as mechanism identity becomes increasingly concentrated rather than distributed throughout the network.
Arm
CE (700.9M)
Δ vs. latin
CE (1.514B)
Δ vs. latin
latin
5.28639
—
5.26963
—
homo_F (homogeneous)
5.37540
+1.68%
5.40807
+2.63%
no_M (drop SSM)
5.39940
+2.14%
5.43837
+3.20%
Table 11: Composition-axis results at two model scales (mean held-out cross-entropy). The same experimental protocol is repeated at 700.9M and 1.514B parameters. The homogeneous-stack penalty and the penalty associated with removing the SSM-family mechanism both increase with scale, indicating that the composition effect is not a small-model artifact.
Task
acc
acc_norm
SciQ
73.7
63.0
PIQA
66.3
65.8
BoolQ
54.3
—
ARC-Easy
52.5
48.7
WinoGrande
51.8
—
HellaSwag
37.2
41.0
Table 12: English benchmark evaluation using lm-evaluation-harness v0.4.11 in the 0-shot setting. Results are reported as task accuracy ( acc ) and normalized accuracy ( acc_norm ) when available.
Task
Score
HellaSwag (acc_norm)
44.6 ± 2.2
COPA
57.2 ± 2.0
SentiNeg
55.7 ± 2.5
WiC
48.8 ± 2.0
BoolQ
47.8 ± 2.0
Table 13: Korean-language evaluation on KoBEST benchmarks. Results are reported as mean task scores with standard deviations across evaluation samples.
Input
Perplexity
Natural Korean sentence
5.4
Same sentence, word order scrambled
41.3
Random tokens
2233.4
Table 14: Perplexity under progressively disrupted linguistic structure. The three inputs contain comparable lexical content but differ substantially in grammatical and compositional organization.
Item
Location
Base model
FINAL-Bench/Aether-7B-5Attn
Instruction-tuned
FINAL-Bench/Aether-7B-5Attn-it
Second checkpoint, same architecture
FINAL-Bench/AETHER-7B-7Attn-base
11-mechanism extension
FINAL-Bench/Aether-6B-11Attn-base
Ablation code
pilot_model.py , pilot_train.py
Causal audit
companion paper, Appendix A
Table 15: Artifacts released with Aether, including model checkpoints, source code, training recipes, and evaluation resources.
Self-attention selects information freely across the sequence, but across depth, Transformers merely add each layer's output to the residual stream, so later layers cannot selectively reuse earlier-layer representations. Recent cross-layer methods improve this flow but operate on hidden states outside attention, adding state beyond the key-value cache at inference--a cost that becomes increasingly salient as modern LLMs compress the cache with grouped-query and multi-head latent attention. We introduce Depth-Attention, which performs this selection inside the attention module itself: before a layer attends over the sequence, its query attends over the keys of earlier layers at the same token position and mixes their values into the value that self-attention then reads. Because Depth-Attention reuses the standard attention queries, keys, and value-cache slots, storing depth-mixed values in place of the original values, it adds no parameters and introduces no persistent inference state beyond the standard key-value cache--the same cache size as a vanilla decoder and less than hidden-state-based cross-layer methods. On Qwen3-style decoders at 1.5B and 3B parameters, Depth-Attention attains the lowest perplexity and the highest average downstream accuracy, improving over the vanilla Transformer by up to 2.3 accuracy points and surpassing strong cross-layer baselines in perplexity and average accuracy, while adding under 0.01% extra arithmetic FLOPs and no additional persistent inference state. The gains hold from 360M to 3B parameters and extend to looped Transformers.
Boyi Zeng, Yiqin Hao, Zitong Wang +7
LUMIA Lab, School of Artificial Intelligence, Shanghai Jiao Tong University · Sun Yat-sen University · Shanghai Innovation Institute +2
Hybrid language models that interleave attention with recurrent components are increasingly competitive with pure Transformers, yet standard LoRA practice applies adapters uniformly without considering the distinct functional roles of each component type. We systematically study component-type LoRA placement across two hybrid architectures -- Qwen3.5-0.8B (sequential, GatedDeltaNet + softmax attention) and Falcon-H1-0.5B (parallel, Mamba-2 SSM + attention) -- fine-tuned on three domains and evaluated on five benchmarks. We find that the attention pathway -- despite being the minority component -- consistently outperforms full-model adaptation with 5-10x fewer trainable parameters. Crucially, adapting the recurrent backbone is destructive in sequential hybrids (-14.8 pp on GSM8K) but constructive in parallel ones (+8.6 pp). We further document a transfer asymmetry: parallel hybrids exhibit positive cross-task transfer while sequential hybrids suffer catastrophic forgetting. These results establish that hybrid topology fundamentally determines adaptation response, and that component-aware LoRA placement is a necessary design dimension for hybrid architectures.
VRAIN – Valencian Research Institute for Artificial Intelligence, Universitat Politècnica de València, Valencia, Spain · Department of Economics and Social Sciences, Universitat Politècnica de València, Valencia, Spain · Department of Business Organisation, Universitat Politècnica de València, Valencia, Spain
Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack of identical layers in which parameters are allocated uniformly across depth. This is a default inherited from the original transformer and largely unchanged since, yet a growing body of evidence suggests that layers contribute non-uniformly to the final output, with later layers refining the residual stream rather than transforming it. We ask whether parameter capacity should reflect this asymmetry. Our controlled experiment shows that, under a fixed budget, allocating more capacity to earlier layers and less to later layers improves perplexity over a uniform-width baseline, while the reverse allocation hurts. Building on this result, we introduce Tapered Language Models (TLMs), an architectural principle in which a parameter-bearing component is monotonically tapered across depth under a fixed total budget. MLPs are the natural site for this instantiation: they dominate parameter count across all modern LM families and expose width as a single, clean axis of variation. Across three model scales and four architectures (Transformer, Gated Attention, Hope-attention, and Titans), tapering MLP width via a smooth cosine schedule consistently improves perplexity and downstream benchmark performance over uniform baselines, at no additional parameter or compute cost. These findings establish depth-aware capacity allocation as a simple, architecture-agnostic axis of language model design, a free lever hidden in plain sight.
Reza Bayat, Ali Behrouz, Aaron Courville
1Mila · 2Cornell University · 3Université de Montréal +1