Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely verified. We formalize prefix invariance and give a lightweight audit, two forward passes, no training or gradients, yielding a per-layer score localizing where causality breaks. Attention-mask inspection, the field's default check, is incomplete: causality is a graph-level property, and leaks can occur via scans, aggregations, or normalization despite correct masks. Across 192 injected-fault trials on eight checkpoints, mask inspection detected none, while our audit localized all 192/192 to the exact layer. Static/dynamic analysis of chunked-scan code in transformers found the same defect in Zamba2 and Nemotron-H, an inter-chunk axis error fixed via the reference implementation. The method fits on one page and runs in seconds.
Figures & tables
Method
Fwd
Bwd
Detected
Exactly localized
False positive on clean model
Ours (per-layer Δ )
2
0
96/96
96/96
none ( Δ=0 )
B1 — logits-only perturbation
2
0
96/96
0/96
none
B2 — static attention-mask inspection
0
0
0/96
0/96
none
B3 — shuffled-suffix perplexity
2
0
71/96
0/96
none
B4 — future-token resampling, 3 seeds
4
0
96/96
0/96
none
B5 — gradient of prefix outputs w.r.t. future positions
1
1
96/96
96/96
none
Table 1: Baseline comparison on 96 injected faults across four architectures. Our method and the gradient-based baseline (B5) both achieve perfect localization, whereas logits-only perturbation (B1) detects leaks but provides no localization. Static attention-mask inspection (B2) detects none of the injected faults.
Model checkpoint
Layers checked
All layers torch.equal
Differing elements
max abs
Bit patterns identical
HuggingFaceTB/SmolLM2-135M
30
[OK]
0
0.0
[OK]
Qwen/Qwen3-0.6B
28
[OK]
0
0.0
[OK]
EleutherAI/pythia-1.4b
24
[OK]
0
0.0
[OK]
state-spaces/mamba-130m-hf
24
[OK]
0
0.0
[OK]
Table 2: Bit-level verification on clean checkpoints. For each checkpoint, all compared layer outputs are exactly identical under repeated evaluation, with zero differing elements and zero maximum absolute difference.
Precision
τ=10−6
τ=10−14
float32
10−6 (L1, L15); 10−7 (L28)
10−10 (all three depths)
float64
10−5 (all three depths)
10−13 (L1, L15); 10−14 (L28)
Table 3: Exact-localization floor: smallest ε still localized to the exact injected layer. Lower values indicate higher sensitivity of the audit to weak temporal-leakage signals.
#
Checkpoint
Mixer class
L
Params
B2
B3
1
HuggingFaceTB/SmolLM2-135M
attention
30
134M
0
18
2
Qwen/Qwen3-0.6B
attention
28
596M
0
17
3
EleutherAI/pythia-1.4b
attention
24
1.41B
0
18
4
google/gemma-3-1b-it
attention
26
1.00B
0
18
5
state-spaces/mamba-130m-hf
state-space (no mask)
24
129M
0
18
6
RWKV/rwkv-6-world-1b6
linear recurrent (no mask)
24
1.60B
0
17
Table 4: Public checkpoint census conducted in float32 on CPU with τ=10−6 and sequence length T=48 . The table summarizes the eight checkpoints audited in the census, together with architecture class, model size, and the outcomes of baselines B2 (static attention-mask inspection) and B3 (shuffled-suffix perplexity).
Checkpoint
Reason
Nature
tiiuae/Falcon-H1-0.5B / 1.5B / 7B
Loaded model does not respond to input; positive control 0/24
Requires a fused normalization kernel from a source build
Custom CUDA kernel dependency
state-spaces/mamba2-2.7b
No model_type in config; not loadable through the standard interface
Packaging
Zyphra/Zamba2-7B
Gated repository; access not granted
Access
Table 5: Attempted but not audited. These checkpoints were examined during the census but could not be evaluated because of loading failures, missing dependencies, packaging issues, or repository-access restrictions.
Check
Result
Negative control (49 layers)
max Δ=0.000e+00
Negative control, after loading trained weights
unchanged 0
Positive control (4 sites × 4 fault classes)
16/16 exact localization
Table 6: Audit of our own released checkpoint (float32, T=48 , τ=10−6 ). The checkpoint passes both negative and positive controls, with exact localization of all injected faults and no leakage observed under normal operation.
Implementation
Inter-chunk axis handling
Class
modeling_mamba2.py (reference)
transpose(1,3) → reduce over input chunk
reference
modeling_bamba.py
matches reference
[OK] conformant
modeling_falcon_h1.py
matches reference
[OK] conformant
modeling_granitemoehybrid.py
matches reference
[OK] conformant
modeling_zamba2.py
permute → reduce over output chunk
departs
modeling_nemotron_h.py
permute → reduce over output chunk
departs
Table 7: Static census of chunked-scan implementations ( transformers 5.7.0). Two implementations ( modeling_zamba2.py and modeling_nemotron_h.py ) depart from the reference inter-chunk recurrence by reducing over the output-chunk axis rather than the input-chunk axis.
Checkpoint
Declared parameter
Verdict
Zamba2-1.2B
chunk_size 256
leak, onset = 256
Nemotron-H-8B
chunk_size 128
leak, onset = 128
Bamba-9B
mamba_chunk_size 256
clean
Falcon-Mamba-7B
conv_kernel 4
clean
Falcon-H1-1.5B
mamba_chunk_size 128
clean
Granite-4.0-H-Tiny
mamba_chunk_size 256
clean (see §3.8.6)
Table 8: Dynamic audit at chunk-exceeding sequence lengths (testing the prediction). Models identified by the static census as deviating from the reference implementation ( Zamba2 and Nemotron-H ) exhibit causal leakage precisely at their declared chunk boundaries, whereas all conformant implementations remain clean.
Alternative explanation
Test
Outcome
Measurement noise
Identical input, two forward passes
Δ=0.000e+00 exactly; computation is deterministic
Hook misuse (shared modules)
Per-layer call counts; module identity
38/38 layers called exactly once; no shared modules
Seed artifact
Four independent random seeds
Reproduced in all four seeds
Checkpoint-specific behavior
Randomly initialized model loaded via from_config
Same onset at 256; structural rather than weight-dependent
Custom-kernel bug
Presence of mamba_ssm or causal_conv1d
Neither installed; the leak appears in the pure-PyTorch implementation
Test-length artifact
Sweep over T=64…512
Clean at T≤257 ; leaks from T≥320 , consistent with a one-chunk degeneracy
Table 9: Alternative-explanation tests for the causal defect observed in Zamba2-1.2B. Each candidate explanation is evaluated through a targeted control experiment. All alternatives are ruled out, indicating that the observed leak is reproducible, structural, and independent of random seeds, trained weights, measurement noise, or implementation-specific kernels.
Model
Maximum prefix Δ
Leak onset
Original
1.12×10−2
256
Patched
0.0000e+00
None
Table 10: Effect of the patch under identical input conditions. The original implementation exhibits measurable prefix contamination with a clear boundary onset, whereas the patched implementation eliminates the effect entirely.
ε
Qwen3-0.6B (control)
Granite-4.0-h
Jamba-tiny
Zamba2-1.2B
10−4
0
∼1.3×10−5
0
2.5×10−5±0.4×10−5
10−3
0
1.34×10−5
0
2.6×10−5±0.8×10−5
10−2
0
1.53×10−5
0
1.2×10−4±1.2×10−4
10−1
0
1.53×10−5
1.48×10−5
1.1×10−3±1.2×10−3
1
0
1.14×10−5
1.53×10−5
7.5×10−3±3.6×10−3
log–log slope
—
≈−0.03 (flat)
≈0 (flat)
0.63±0.07
Table 11: Perturbation sweep of prefix-delta response across multiple architectures. Values report the measured prefix deviation as a function of perturbation magnitude ε . The Zamba2-1.2B column reports the mean and standard deviation over four random seeds, while the remaining models are deterministic single-run measurements. A non-zero log–log slope indicates amplification of future perturbations into the prefix.
Class
Pattern
Transformation
C1
shift
concat(o[:, 1:], o[:, -1:]) — output pulled one step earlier
C1
peek1
o + eps * concat(o[:, 1:], o[:, -1:])
C2
pool_leak
o + eps * mean(o, dim=time, keepdim)
C2
global_max
o + eps * max(o, dim=time, keepdim)
C2
future_mean_k
o + eps * mean of the next k = 4 positions
C3
norm_leak
(1-eps)o + eps(o-mu)/sigma , with μ , σ over (time, hidden) jointly
Table 12: Executable injected-fault specifications grouped into four classes of temporal leakage. Each transformation is applied to a selected layer output and serves as a positive-control fault throughout the evaluation.
File
Role
leak_cpu_suite2.py
Census harness: detector + injections + B1–B3; per-model JSON flush; records failures with traceback
Table 13: Primary experimental scripts used throughout the evaluation. The table summarizes the main audit harnesses, baseline implementations, and sensitivity-analysis programs used to generate the reported results.
File
Contents
results/smoke_smol.json
SmolLM2-135M census record
results/tier1.json , results/tier1__*.json
Qwen3-0.6B, mamba-130m-hf
results/b3__LiquidAI_LFM2-1.2B.json
LFM2-1.2B, including the §3.5 floor anomaly ( sensitivity.sweep )
results/b5__RWKV_rwkv-6-world-1b6.json
RWKV-6 1.6B
results/b5__google_gemma-3-1b-it.json
Gemma-3-1B-it
results/b5__google_recurrentgemma-2b.json
RecurrentGemma-2B
Table 14: Result files associated with the public-checkpoint census (§3.6) and the attempted-but-not-audited cases summarized in Tables 4 and 5. The table maps released JSON artifacts to the models, audits, and failure analyses discussed in the paper.
File
Contents
results/e2_smol_fp32.json
SmolLM2-135M: all six baselines including B6; clean_bitwise ; clean-model B6 false positive 1.850128173828125e-04
results/e2_multi_nob6.json
Qwen3-0.6B, pythia-1.4b, mamba-130m-hf: B1–B5; clean_bitwise for each
Table 15: Result files associated with the baseline comparison (§3.3, Table 1) and bit-level verification experiments (§3.4, Table 2). The table maps released JSON artifacts to the corresponding evaluations and checkpoints.
File
Contents
results/sens_smol_fp32.json
float32, τ=10−6 , 45 points (3 depths × 15 ε )
results/sens_smol_fp64_v2.json
float64, τ=10−6 , 45 points
results/sens_smol_fp32_tau1e-14.json
float32, τ=10−14
results/sens_smol_fp64_tau1e-14.json
float64, τ=10−14
results/sens_smol_fp64.json
SUPERSEDED — do not cite. Contains the Appendix A.5 float32-cast defect
Table 16: Result files associated with the precision and threshold sensitivity analysis (§3.5, Table 3). The table lists the released JSON artifacts used to characterize exact-localization floors under different numerical precisions and threshold settings.