Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems round the state stochastically; we ask which rounding rule such caches should use. We find that a deterministic golden-ratio Weyl dither, which needs no random numbers, consistently brings the quantized model closer to the full-precision one than stochastic rounding, across pure and hybrid models, storage formats, and long decoding horizons, at no extra cost. Round-to-nearest behaves differently: because it discards small updates, its error keeps growing, so it can look best in short evaluations yet falls far behind over long generations. A discrepancy analysis explains this ordering, and we document implementation pitfalls that silently remove the benefit.
Figures & tables
INT8
FP8 (E4M3)
BF16
Model
rtn
sr
Weyl
rtn
sr
Weyl
rtn
sr
Weyl
Mamba-1 130M
1.73
0.88
0.58
15.7
8.35
6.40
1.98
0.215
0.129
Mamba-1 370M
1.75
0.65
0.44
17.0
6.48
4.98
1.64
0.157
0.095
Mamba-2 130M
26.9
4.15
2.83
23.9
23.5
17.6
3.55
0.707
0.358
Granite 4.0-H 350M
7.69
22.7
16.3
41.6
114
68.3
1.15
1.74
1.24
Granite 4.0-H 1B
3.12
5.55
3.94
5.70
12.7
11.3
0.426
0.542
0.375
Table 1: Decode regime: KL divergence to the full-precision model ( ×10−3 , lower is better) for the three rounding rules, and the relative KL reduction of the Weyl dither over sr , estimate [95% lower bound]. INT8 uses an FP32 block scale. Pure Mamba: 32 WikiText-103 chunks, 1024 quantized positions after a 1024-token full-precision prefix, sr over two seeds. Hybrids: 32 PG-19 books, 512 quantized decoding steps after a 2048-token prefill, sr with one seed. Bold: best rule in the row.
Figure 1: How the damage evolves over a long generation. KL per token, averaged over consecutive blocks of 512 quantized decoding steps, for three hybrid models (columns) and two storage formats (rows; INT8 with an FP32 block scale); 8 PG-19 books, 2048-token full-precision prefill. rtn ’s error keeps growing, while the two dithers level off; the Weyl dither stays below sr throughout.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Mamba-1 130M, every-token regime. (a) RMS stored-state error in quantizer steps (INT8) by memory length: rtn is smallest. (b) Output KL to FP32 at 2k tokens relative to sr : rtn is worst.
KL ( ×10−3 )
Model
Fmt
rtn
sr
Weyl
Weyl vs. sr (%)
130M
FP8
16.65
1.81
1.28
29.4 [12.2, 44.8]
130M
INT8
4.22
0.24
0.19
21.5 [ − 4.1, 38.2]
130M
INT4
142.6
20.7
13.7
33.9 [10.5, 51.5]
370M
FP8
5.60
0.86
0.54
36.8 [20.6, 49.5]
370M
INT8
1.89
0.09
0.05
†
Appendix
Table 2: Free-running generation (1024 greedy tokens, 64 prompts): per-step KL between the quantized-cache model and an FP32 cache fed the same tokens; relative reduction of Weyl vs. sr with 95% CI over prompts. † Below resolution ( sr KL <10−4 ).
Model
FP32
rtn
sr
Weyl
Weyl − sr
Mamba-1 130M
44.25
37.92
40.81
43.00
+2.19
Mamba-1 370M
55.62
46.46
52.84
53.66
+0.82
Mamba-1 790M
61.71
55.37
59.93
60.37
+0.45
Mamba-1 1.4B
64.95
58.47
63.21
63.44
+0.23
Mamba-2 130M
43.97
–
38.54
39.22
+0.68
Mamba-2 370M
55.93
–
51.27
52.82
+1.55
Appendix
Table 3: LAMBADA accuracy (%) with an INT4 state cache (every-token regime).
Model
Format
rtn
sr
Hash
Weyl
Opt.
Gated
350M
FP16 †
0.071
0.048
–
0.032
–
–
350M
BF16
1.16
1.68
–
1.29
–
–
350M
FP8
43.7
107
–
64.8
–
–
350M
INT8 (128)
7.98
24.7
22.6
17.4
16.6
15.4
350M
INT8 (32)
5.66
13.4
–
10.5
–
–
350M
INT8 (16)
4.07
8.98
–
7.00
–
6.00
Appendix
Table 4: Granite, decode regime, all rules at 2K prefill (KL ×10−3 ). Hash: integer-hash sr ; Opt.: Weyl with offsets optimized for neighbor separation; Gated: rtn for entries whose magnitude is below one step, Weyl above (exploratory). † Below the floor ( 10−4 ).
Model
Format
Measured
Replay
Composed
Sw
Mamba-1 130M
INT8
2.53
1.55
1.75
0.83
FP8
2.44
1.73
1.90
0.89
BF16
15.80
10.26
11.78
0.62
Mamba-1 370M
INT8
3.75
2.06
3.17
0.83
FP8
3.43
1.90
2.43
0.90
BF16
17.86
16.17
13.44
0.62
Appendix
Table 5: The rtn reversal, INT8, FP8, and BF16. Measured: KL ratio rtn /Weyl (Tables 13 , 4 ). Replay: state-error ratio from the open-loop replay. Composed: replay errors weighted by the measured noise sensitivities. Sw : memory-weighted fraction of writes smaller than half a step. Ratios >1 mean the dither wins.
Replay rtn /Weyl by bin
Sensitivity by bin
Mamba-1 130M
0.51
2.00
1.91
4.41
0.013
0.065
0.361
0.021
Mamba-1 370M
1.16
2.62
3.21
3.98
0.047
0.121
0.381
0.028
Mamba-2 130M
0.69
1.71
2.67
1.45
0.150
1.124
0.390
0.045
Granite 350M
0.69
1.54
1.08
1.29
44.6
1.21
13.6
0.61
Granite 1B
0.74
1.49
1.22
0.62
3.93
0.88
1.33
0.049
Appendix
Table 6: Replay by memory bin (INT8, primary block) and noise sensitivity (KL per unit state error) by memory bin τ<10 , 10–100, 100–1000, ≥1000 .
Mamba-1 130M
Mamba-1 370M
Format
Rotation
rtn
sr
Weyl
rtn
sr
Weyl
FP16
–
0.00007
0.00001
0.00000
0.00008
0.00000
0.00000
BF16
–
0.00372
0.00027
0.00016
0.00372
0.00020
0.00012
FP8
–
0.02751
0.01002
0.00761
0.03332
0.00832
0.00628
INT8
–
0.00249
0.00122
0.00091
0.00322
0.00095
0.00071
INT4
–
0.17835
0.13332
0.10471
0.15621
0.09988
0.07931
Appendix
Table 7: All formats, Mamba-1 at 2k tokens (KL to FP32; sr and Weyl averaged over two seeds), including block-16 Hadamard rotation over the state dimension (these rotation runs predate the quantizer correction of Appendix E ; the corrected INT4 values are in Section 6.5 ). FP16 and BF16 rows, and INT8 for 370M, are below the resolution floor ( sr KL <10−3 ) and are not used for any claim.
Dither
LD256
130M INT4
130M FP8
370M INT4
370M FP8
Weyl, α=2−1
0.0069
0.10413
0.00761
0.07783
0.00625
Weyl, α=1/(2+φ)
0.0075
0.10429
0.00780
0.08086
0.00624
Weyl, α=φ−1
0.0078
0.10499
0.00763
0.07999
0.00619
Weyl, α=e−2
0.0097
0.10606
0.00789
0.08042
0.00648
Weyl, per-entry α∈[0.2,0.8]
0.0188
0.12100
0.00892
0.09326
0.00737
sr
0.0539
0.13242
0.00996
0.09970
0.00825
Appendix
Table 8: Dither sequences at 2k tokens (KL to FP32). LD256 : local star discrepancy.
FP16 scale
FP32 scale
Weyl+FP32 vs.
Model
sr
Weyl
sr
Weyl
sr +FP16
Mamba-1 130M
0.00120
0.00091
0.00097
0.00065
−45.8%
Mamba-1 370M
0.00096
0.00072
0.00079
0.00053
−44.8%
Mamba-1 790M
0.00073
0.00054
0.00061
0.00041
−43.8%
Mamba-1 1.4B
0.00083
0.00059
0.00072
0.00044
−47.0%
Appendix
Table 9: FP16 versus FP32 block scale at INT8 (KL to FP32, 2k tokens). Rounding the FP16 scale to nearest clips the block maximum in 50.1% of blocks (unit test on 105 random blocks). At INT4 the scale precision has no measurable effect (Mamba-1 130M Weyl: 0.10499 vs. 0.10509; sr : 0.13242 vs. 0.13111).
Model
Format
Weyl (float)
Knuth (fixed pt.)
Non-aliased (fixed pt.)
Aliased (float)
Mamba-1 130M
FP8
0.00806
0.00787
0.00790
–
Mamba-1 130M
INT4
0.11161
0.11159
0.11118
–
Mamba-2 130M
FP8
0.01880
0.01954
0.01894
0.01929
Mamba-2 130M
INT8
0.00323
0.00325
0.00322
0.00325
Mamba-2 130M
INT4
0.53080
0.74169
0.50917
0.70346
Mamba-2 370M
FP8
0.00747
0.00846
0.00737
–
Appendix
Table 10: Offsets for the hardware dither (KL to FP32, 2k tokens). Knuth: offsets from the multiplicative hash (aliased with the time increment). Non-aliased: independent generators. Aliased float: ui,t={(i+t)φ} . All rows use 16 evaluation chunks (the Mamba-1 130M entries in Table 12 use 32).
Values
copy
rtn
sr
Weyl (Knuth)
Weyl (non-aliased)
sr / rtn
Weyl/ rtn
1.0M
17.7
25.2
49.6
25.4
25.2
1.97
1.00
4.2M
40.5
43.5
168.8
48.7
49.2
3.88
1.13
16.8M
153.5
154.7
472.6
167.5
166.3
3.06
1.08
Appendix
Table 11: State-write kernel on a T4 (median microseconds over 7 interleaved rounds of 50 launches; INT8 storage, 16 values per FP16 scale, synthetic in-kernel update).
FP8 (10 bits)
INT8 (9 bits)
INT4 (5 bits)
Model
PPL
rtn
sr
Weyl
rtn
sr
Weyl
rtn
sr
Weyl
Mamba-1 130M
21.64
.0275
.0100
.0076
.0025
.0012
.0009
.178
.132
.105
Mamba-1 370M
15.09
.0333
.0083
.0062
.0032
.0010 †
.0007
.156
.100
.080
Mamba-1 790M
12.54
–
.0064
.0048
–
.0007 †
.0005
–
.069
.053
Mamba-1 1.4B
11.29
–
.0074
.0053
–
.0008 †
.0006
–
.075
.060
Mamba-2 130M
21.24
.1271
.0263
.0188
.0538
.0046
.0032
1.614
.679
.531
Appendix
Table 12: Every-token regime (WikiText-103, 2048 tokens): KL to the FP32 model, and the relative KL reduction of the Weyl dither over sr with 95% paired-bootstrap intervals. † sr KL below the resolution floor of 10−3 ; reported, not counted. rtn was not run for 790M and 1.4B.
BF16
FP8
INT8
INT4
Model
rtn
sr
Weyl
rtn
sr
Weyl
rtn
sr
Weyl
rtn
sr
Weyl
Mamba-1 130M
2.07
0.21
0.13
15.6
8.37
6.41
1.90
1.03
0.75
88.7
95.4
77.7
Mamba-1 370M
1.74
0.17
0.10
17.4
6.56
5.07
2.13
0.79
0.57
78.3
67.0
54.8
Mamba-2 130M
3.63
0.66
0.34
22.5
21.3
16.7
22.8
4.16
2.92
178
506
421
Granite 350M
1.16
1.68
1.29
43.7
107
64.8
7.98
24.7
17.4
–
449
402
Granite 1B
0.44
0.53
0.37
5.64
12.5
10.4
3.22
6.05
4.49
–
137
112
Appendix
Table 13: Decode regime (FP32 prefill, quantized decoding): KL to the FP32 model ( ×10−3 ) and the Weyl-vs- sr KL reduction, estimate [95% lower bound] in %. Mamba: WikiText-103, 1024-token prefill, 1024 decoding positions, sr and Weyl averaged over 5 (130M) or 3 seeds. Granite: PG-19, 2048-token prefill, 512 decoding steps. † Below the resolution floor; not counted. rtn INT4 was not run on Granite.
KL, block 128
Weyl vs. sr (%)
Model
Prefill
rtn
sr
Weyl
128
16
CC=8
350M
2K
7.98
24.7
17.4
29 [21]
22 [15]
− 2 [ − 18]
350M
8K
7.87
24.6
16.6
33 [24]
24 [17]
9 [ − 3]
350M
32K
9.71
31.8
19.0
40 [29]
24 [6]
31 [12]
1B
2K
3.22
6.05
4.49
26 [21]
20 [16]
1 [ − 5]
1B
8K
2.77
5.79
4.76
18 [10]
26 [20]
4 [ − 12]
Appendix
Table 14: Granite, decode regime, INT8: KL ( ×10−3 , block 128) by prefill length, and the Weyl-vs- sr reduction, estimate [95% lower bound] in %, at block 128, block 16, and with checkpointing (CC=8: FP32 intermediate states, rounding every 8th write; block 128). 16 documents, except 8 (1B, 8K; 350M, 32K) and 4 (1B, 32K).
Figure 3: (a) Output sensitivity to noise injected into one memory bin, per unit of state error, relative to the short-memory bin ( τ<10 ). (b) Measured KL ratio rtn /Weyl versus the prediction from state error alone (open) and from state error weighted by the measured sensitivities (filled); 19 cells, INT8, FP8, and BF16. Granite cells in the upper-left quadrant are mispredicted.
First 512
Last 512
Weyl vs. sr (%)
Model
Fmt
rtn
sr
Weyl
rtn
sr
Weyl
First
Last
Granite 350M
INT8
8.36
23.0
16.7
32.7
52.4
38.0
27.5 [22.0]
27.5 [23.8]
Granite 350M
BF16
1.21
1.69
1.36
22.2
4.22
2.90
19.3 [ − 0.7]
31.3 [13.3]
Granite 1B
INT8
3.08
5.42
3.95
12.2
12.1
9.63
27.2 [19.8]
20.7 [13.2]
Granite 1B
BF16
0.446
0.559
0.372
9.61
1.94
1.37
33.5 [26.4]
29.4 [19.1]
Falcon-H1 0.5B
INT8
2.47
1.41
1.43
92.9
2.60
1.96
− 1.3 [ − 39.8]
24.6 [13.3]
Appendix
Table 15: KL per token ( ×10−3 ) in the first and last 512-step blocks of 4096 quantized decoding steps, and the Weyl-vs- sr reduction in those blocks, estimate [95% lower bound]; 8 documents.