Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely verified. We formalize prefix invariance and give a lightweight audit, two forward passes, no training or gradients, yielding a per-layer score localizing where causality breaks. Attention-mask inspection, the field's default check, is incomplete: causality is a graph-level property, and leaks can occur via scans, aggregations, or normalization despite correct masks. Across 192 injected-fault trials on eight checkpoints, mask inspection detected none, while our audit localized all 192/192 to the exact layer. Static/dynamic analysis of chunked-scan code in transformers found the same defect in Zamba2 and Nemotron-H, an inter-chunk axis error fixed via the reference implementation. The method fits on one page and runs in seconds.
Figures & tables
Method
Fwd
Bwd
Detected
Exactly localized
False positive on clean model
Ours (per-layer Δ )
2
0
96/96
96/96
none ( Δ=0 )
B1 — logits-only perturbation
2
0
96/96
0/96
none
B2 — static attention-mask inspection
0
0
0/96
0/96
none
B3 — shuffled-suffix perplexity
2
0
71/96
0/96
none
B4 — future-token resampling, 3 seeds
4
0
96/96
0/96
none
B5 — gradient of prefix outputs w.r.t. future positions
1
1
96/96
96/96
none
Table 1: Baseline comparison on 96 injected faults across four architectures. Our method and the gradient-based baseline (B5) both achieve perfect localization, whereas logits-only perturbation (B1) detects leaks but provides no localization. Static attention-mask inspection (B2) detects none of the injected faults.
Model checkpoint
Layers checked
All layers torch.equal
Differing elements
max abs
Bit patterns identical
HuggingFaceTB/SmolLM2-135M
30
[OK]
0
0.0
[OK]
Qwen/Qwen3-0.6B
28
[OK]
0
0.0
[OK]
EleutherAI/pythia-1.4b
24
[OK]
0
0.0
[OK]
state-spaces/mamba-130m-hf
24
[OK]
0
0.0
[OK]
Table 2: Bit-level verification on clean checkpoints. For each checkpoint, all compared layer outputs are exactly identical under repeated evaluation, with zero differing elements and zero maximum absolute difference.
Precision
τ=10−6
τ=10−14
float32
10−6 (L1, L15); 10−7 (L28)
10−10 (all three depths)
float64
10−5 (all three depths)
10−13 (L1, L15); 10−14 (L28)
Table 3: Exact-localization floor: smallest ε still localized to the exact injected layer. Lower values indicate higher sensitivity of the audit to weak temporal-leakage signals.
#
Checkpoint
Mixer class
L
Params
B2
B3
1
HuggingFaceTB/SmolLM2-135M
attention
30
134M
0
18
2
Qwen/Qwen3-0.6B
attention
28
596M
0
17
3
EleutherAI/pythia-1.4b
attention
24
1.41B
0
18
4
google/gemma-3-1b-it
attention
26
1.00B
0
18
5
state-spaces/mamba-130m-hf
state-space (no mask)
24
129M
0
18
6
RWKV/rwkv-6-world-1b6
linear recurrent (no mask)
24
1.60B
0
17
Table 4: Public checkpoint census conducted in float32 on CPU with τ=10−6 and sequence length T=48 . The table summarizes the eight checkpoints audited in the census, together with architecture class, model size, and the outcomes of baselines B2 (static attention-mask inspection) and B3 (shuffled-suffix perplexity).
Checkpoint
Reason
Nature
tiiuae/Falcon-H1-0.5B / 1.5B / 7B
Loaded model does not respond to input; positive control 0/24
Requires a fused normalization kernel from a source build
Custom CUDA kernel dependency
state-spaces/mamba2-2.7b
No model_type in config; not loadable through the standard interface
Packaging
Zyphra/Zamba2-7B
Gated repository; access not granted
Access
Table 5: Attempted but not audited. These checkpoints were examined during the census but could not be evaluated because of loading failures, missing dependencies, packaging issues, or repository-access restrictions.
Check
Result
Negative control (49 layers)
max Δ=0.000e+00
Negative control, after loading trained weights
unchanged 0
Positive control (4 sites × 4 fault classes)
16/16 exact localization
Table 6: Audit of our own released checkpoint (float32, T=48 , τ=10−6 ). The checkpoint passes both negative and positive controls, with exact localization of all injected faults and no leakage observed under normal operation.
Implementation
Inter-chunk axis handling
Class
modeling_mamba2.py (reference)
transpose(1,3) → reduce over input chunk
reference
modeling_bamba.py
matches reference
[OK] conformant
modeling_falcon_h1.py
matches reference
[OK] conformant
modeling_granitemoehybrid.py
matches reference
[OK] conformant
modeling_zamba2.py
permute → reduce over output chunk
departs
modeling_nemotron_h.py
permute → reduce over output chunk
departs
Table 7: Static census of chunked-scan implementations ( transformers 5.7.0). Two implementations ( modeling_zamba2.py and modeling_nemotron_h.py ) depart from the reference inter-chunk recurrence by reducing over the output-chunk axis rather than the input-chunk axis.
Checkpoint
Declared parameter
Verdict
Zamba2-1.2B
chunk_size 256
leak, onset = 256
Nemotron-H-8B
chunk_size 128
leak, onset = 128
Bamba-9B
mamba_chunk_size 256
clean
Falcon-Mamba-7B
conv_kernel 4
clean
Falcon-H1-1.5B
mamba_chunk_size 128
clean
Granite-4.0-H-Tiny
mamba_chunk_size 256
clean (see §3.8.6)
Table 8: Dynamic audit at chunk-exceeding sequence lengths (testing the prediction). Models identified by the static census as deviating from the reference implementation ( Zamba2 and Nemotron-H ) exhibit causal leakage precisely at their declared chunk boundaries, whereas all conformant implementations remain clean.
Alternative explanation
Test
Outcome
Measurement noise
Identical input, two forward passes
Δ=0.000e+00 exactly; computation is deterministic
Hook misuse (shared modules)
Per-layer call counts; module identity
38/38 layers called exactly once; no shared modules
Seed artifact
Four independent random seeds
Reproduced in all four seeds
Checkpoint-specific behavior
Randomly initialized model loaded via from_config
Same onset at 256; structural rather than weight-dependent
Custom-kernel bug
Presence of mamba_ssm or causal_conv1d
Neither installed; the leak appears in the pure-PyTorch implementation
Test-length artifact
Sweep over T=64…512
Clean at T≤257 ; leaks from T≥320 , consistent with a one-chunk degeneracy
Table 9: Alternative-explanation tests for the causal defect observed in Zamba2-1.2B. Each candidate explanation is evaluated through a targeted control experiment. All alternatives are ruled out, indicating that the observed leak is reproducible, structural, and independent of random seeds, trained weights, measurement noise, or implementation-specific kernels.
Model
Maximum prefix Δ
Leak onset
Original
1.12×10−2
256
Patched
0.0000e+00
None
Table 10: Effect of the patch under identical input conditions. The original implementation exhibits measurable prefix contamination with a clear boundary onset, whereas the patched implementation eliminates the effect entirely.
ε
Qwen3-0.6B (control)
Granite-4.0-h
Jamba-tiny
Zamba2-1.2B
10−4
0
∼1.3×10−5
0
2.5×10−5±0.4×10−5
10−3
0
1.34×10−5
0
2.6×10−5±0.8×10−5
10−2
0
1.53×10−5
0
1.2×10−4±1.2×10−4
10−1
0
1.53×10−5
1.48×10−5
1.1×10−3±1.2×10−3
1
0
1.14×10−5
1.53×10−5
7.5×10−3±3.6×10−3
log–log slope
—
≈−0.03 (flat)
≈0 (flat)
0.63±0.07
Table 11: Perturbation sweep of prefix-delta response across multiple architectures. Values report the measured prefix deviation as a function of perturbation magnitude ε . The Zamba2-1.2B column reports the mean and standard deviation over four random seeds, while the remaining models are deterministic single-run measurements. A non-zero log–log slope indicates amplification of future perturbations into the prefix.
Class
Pattern
Transformation
C1
shift
concat(o[:, 1:], o[:, -1:]) — output pulled one step earlier
C1
peek1
o + eps * concat(o[:, 1:], o[:, -1:])
C2
pool_leak
o + eps * mean(o, dim=time, keepdim)
C2
global_max
o + eps * max(o, dim=time, keepdim)
C2
future_mean_k
o + eps * mean of the next k = 4 positions
C3
norm_leak
(1-eps)o + eps(o-mu)/sigma , with μ , σ over (time, hidden) jointly
Table 12: Executable injected-fault specifications grouped into four classes of temporal leakage. Each transformation is applied to a selected layer output and serves as a positive-control fault throughout the evaluation.
File
Role
leak_cpu_suite2.py
Census harness: detector + injections + B1–B3; per-model JSON flush; records failures with traceback
Table 13: Primary experimental scripts used throughout the evaluation. The table summarizes the main audit harnesses, baseline implementations, and sensitivity-analysis programs used to generate the reported results.
File
Contents
results/smoke_smol.json
SmolLM2-135M census record
results/tier1.json , results/tier1__*.json
Qwen3-0.6B, mamba-130m-hf
results/b3__LiquidAI_LFM2-1.2B.json
LFM2-1.2B, including the §3.5 floor anomaly ( sensitivity.sweep )
results/b5__RWKV_rwkv-6-world-1b6.json
RWKV-6 1.6B
results/b5__google_gemma-3-1b-it.json
Gemma-3-1B-it
results/b5__google_recurrentgemma-2b.json
RecurrentGemma-2B
Table 14: Result files associated with the public-checkpoint census (§3.6) and the attempted-but-not-audited cases summarized in Tables 4 and 5. The table maps released JSON artifacts to the models, audits, and failure analyses discussed in the paper.
File
Contents
results/e2_smol_fp32.json
SmolLM2-135M: all six baselines including B6; clean_bitwise ; clean-model B6 false positive 1.850128173828125e-04
results/e2_multi_nob6.json
Qwen3-0.6B, pythia-1.4b, mamba-130m-hf: B1–B5; clean_bitwise for each
Table 15: Result files associated with the baseline comparison (§3.3, Table 1) and bit-level verification experiments (§3.4, Table 2). The table maps released JSON artifacts to the corresponding evaluations and checkpoints.
File
Contents
results/sens_smol_fp32.json
float32, τ=10−6 , 45 points (3 depths × 15 ε )
results/sens_smol_fp64_v2.json
float64, τ=10−6 , 45 points
results/sens_smol_fp32_tau1e-14.json
float32, τ=10−14
results/sens_smol_fp64_tau1e-14.json
float64, τ=10−14
results/sens_smol_fp64.json
SUPERSEDED — do not cite. Contains the Appendix A.5 float32-cast defect
Table 16: Result files associated with the precision and threshold sensitivity analysis (§3.5, Table 3). The table lists the released JSON artifacts used to characterize exact-localization floors under different numerical precisions and threshold settings.
What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length? Existing answers often rely on idealized arithmetic, but under finite precision, rounding and evaluation order can change what information attention retains and therefore what the model can compute. We develop an algebraic formalization that derives expressivity directly from the model's implemented dynamics. Its central object is its memory; the finite internal state computed by attention that summarizes the information from the prefix available to all future queries. Each attention head updates its own state independently within a layer, while layers compose hierarchically, providing a uniform route from model assumptions to expressivity bounds. Applying this method to transformers without positional embeddings, we obtain an expressivity hierarchy governed by the attention type under specific numerical semantics. Width-one sliding-window attention supports bounded-suffix memory, while a modified form of soft attention supports irreversible, checklist-like state, and combining the two mechanisms provides an interplay of both. Ordinary left-to-right floating-point soft attention can realize more expressive memory operations than any of the above. Algebraically, the four cases correspond to definite, R-trivial, locally R-trivial, and aperiodic semigroups. Under an explicit free-wiring assumption, all four bounds are tight.
Probes are routinely paired with an intervention: ablate the direction the probe found, run the model, and read the change in task accuracy, taking a large drop as evidence that the computation depends on what the probe read and a near-zero drop as evidence that it does not. Either inference requires that the ablation have removed the target from the layer. We find that the ablation does not remove what it targets. A probe refitted on the ablated activations recovers its original accuracy in every cell we test, and keeps recovering when the probe's entire row space is deleted rather than a single axis, because the quantity survives in the orthogonal complement. Because a refitted probe recovers, neither a large task drop nor a near-zero one establishes whether the model needed the target, and one probe fit detects this. Replacing the ablation with iterative nullspace projection, scored against random subspaces of matched dimension, reverses the conclusion: representations that looked causally inert carry most of the task. The correction also separates where a variable is most readable from where deleting it does most damage, and those are not the same layer in any pretrained model we study. The erasure is defined by a linear probe family, so removing a nonlinearly encoded quantity remains open.
Transformers do not merely lack data on some Boolean extrapolation tasks; they generalize in a systematically wrong way. Recent work on generalization on the unseen has shown that, despite fitting the observed domain, Transformers often extrapolate according to a simpler minimum-degree interpolator rather than the true target function. These Boolean tasks are not practical applications, but controlled stress tests for understanding Transformer inductive bias. We ask whether this failure mode can be corrected by injecting explicit structural priors into attention. Existing structured-initialization methods alter Transformer inductive bias indirectly, by choosing query and key projections whose similarity scores approximate a desired attention pattern. However, we find that when applied to Boolean extrapolation, these QK-based priors can be rapidly overwritten during training and fail to change the learned extrapolation rule. We propose a simpler alternative: initialize the additive attention mask directly. Unlike standard hard masks used for causality or locality attention, our mask is a finite, learnable attention-logit bias initialized from task-level interaction structure. This separates the structural prior from content-dependent attention scores, allowing it to persist throughout optimization. On Boolean reasoning tasks, mask-based initialization achieves near-perfect extrapolation where vanilla and QK-initialized Transformers remain trapped by the default inductive bias. The same mechanism also improves low-data arithmetic performance and remains competitive on vision and language benchmarks. These results show that attention masks can serve not only as architectural constraints, but as a simple substrate for encoding persistent inductive bias in Transformers.
Mingze Ma, Hemanth Saratchandran, Cameron Gordon +1
Australian Institute for Machine Learning · Adelaide University