While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon. We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention. Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats. We propose the specifications for a dedicated accelerator capable of executing this pipeline efficiently. Simulated experiments with Qwen3 and Gemma 3 show that rerouting a selective minority of sub-blocks to high precision is sufficient to recover the baseline model performance.
Figures & tables
Perplexity (↓)
16-bit sub-blocks per tile
Model name
8-bit
8/16-bit LAMP
32-bit
0
1
2+
gemma-3-27b-pt
21.470
17.150
16.540
46.30%
20.52%
33.18%
gemma-3-12b-pt
36.740
21.830
18.680
50.94%
20.30%
28.76%
Qwen3-32B
28.180
23.740
23.800
65.09%
15.43%
19.48%
Qwen3-8B
44.030
33.780
33.340
66.19%
14.58%
19.23%
Qwen3-30B-A3B
43.900
28.270
27.510
69.75%
13.71%
16.54%
Table 1: LampAttention with δ=2−8 and τ=2−1 on C4.
Accuracy (↑)
16-bit sub-blocks per tile
Model name
8-bit
8/16-bit LAMP
32-bit
0
1
2+
gemma-3-27b-pt
0.6261
0.7851
0.7993
48.49%
21.55%
29.96%
gemma-3-12b-pt
0.3618
0.5899
0.7654
53.39%
22.19%
24.42%
Qwen3-32B
0.7555
0.8355
0.8344
54.52%
19.67%
25.81%
Qwen3-8B
0.5625
0.7412
0.7840
54.05%
19.11%
26.84%
Qwen3-30B-A3B
0.5998
0.7884
0.8169
56.97%
18.99%
24.04%
Table 2: LampAttention with δ=2−8 and τ=2−1 on 5-shot MMLU.
Figure 1: LampAttention with δ=2−8 and τ∈{2−t:t=0,…,5} .
Figure 2: LampAttention with δ=2−8 and τ∈{2−t:t=0,…,5} .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Perplexity (↓)
16-bit sub-blocks per tile
Model name
8-bit
8/16-bit LAMP
32-bit
0
1
2+
gemma-3-27b-pt
8.6290
6.2100
5.9000
47.21%
19.86%
32.93%
gemma-3-12b-pt
15.940
8.6370
7.3790
52.86%
19.10%
28.04%
Qwen3-32B
12.170
9.4290
9.3610
76.50%
11.55%
11.95%
Qwen3-8B
17.480
12.510
12.300
76.44%
11.01%
12.55%
Qwen3-30B-A3B
18.780
11.260
10.890
81.82%
9.02%
9.16%
Appendix
Table 3: LampAttention with δ=2−8 and τ=2−1 on Wikitext.
Accuracy (↑)
16-bit sub-blocks per tile
Model name
8-bit
8/16-bit LAMP
32-bit
0
1
2+
gemma-3-27b-pt
0.5977
0.6484
0.6562
48.47%
30.41%
21.12%
gemma-3-12b-pt
0.4570
0.6191
0.6387
48.71%
32.21%
19.08%
Qwen3-32B
0.6191
0.6016
0.6094
48.54%
31.22%
20.24%
Qwen3-8B
0.4766
0.5352
0.5605
48.24%
28.71%
23.05%
Qwen3-30B-A3B
0.4512
0.5176
0.5762
48.28%
28.94%
22.78%
Appendix
Table 4: LampAttention with δ=2−8 and τ=2−1 on 0-shot ARC-Challenge.
Figure 3: LampAttention with δ=2−8 and τ∈{2−t:t=0,…,5} .
Figure 4: LampAttention with δ∈{2−8,2−4} and τ∈{2−t:t=0,…,5} .
Figure 5: LampAttention with δ∈{2−8,2−4} and τ∈{2−t:t=0,…,5} .
Figure 6: LampAttention with δ=2−8 or disabled first stage, and τ∈{2−t:t=0,…,5} .
Figure 7: LampAttention with δ=2−8 or disabled first stage, and τ∈{2−t:t=0,…,5} .
Accuracy (↑)
16-bit sub-blocks per tile
Model name
8-bit
8/16-bit LAMP
32-bit
0
1
2+
gemma-3-27b-pt
0.6504
0.6953
0.7070
56.64%
19.37%
23.99%
gemma-3-12b-pt
0.4609
0.6289
0.6777
62.56%
19.31%
18.13%
Qwen3-32B
0.6758
0.7305
0.7324
55.23%
19.55%
25.22%
Qwen3-8B
0.5488
0.6660
0.6660
54.10%
18.53%
27.37%
Qwen3-30B-A3B
0.5879
0.7031
0.6953
57.72%
19.01%
23.27%
Appendix
Table 5: LampAttention with δ=2−8 and τ=2−1 on 25-shot ARC-Challenge.
Accuracy (↑)
16-bit sub-blocks per tile
Model name
8-bit
8/16-bit LAMP
32-bit
0
1
2+
gemma-3-27b-pt
0.5977
0.6484
0.6562
48.47%
30.41%
21.12%
gemma-3-12b-pt
0.4570
0.6191
0.6387
48.71%
32.21%
19.08%
Qwen3-32B
0.6191
0.6016
0.6094
48.54%
31.22%
20.24%
Qwen3-8B
0.4766
0.5352
0.5605
48.24%
28.71%
23.05%
Qwen3-30B-A3B
0.4512
0.5176
0.5762
48.28%
28.94%
22.78%
Appendix
Table 6: LampAttention with δ=2−8 and τ=2−1 on 0-shot ARC-Challenge.
Figure 8: LampAttention with δ=2−8 and τ∈{2−t:t=0,…,5} .
Modern large language models increasingly require long contexts for reasoning and multi-document tasks, but attention's quadratic complexity creates a severe computational bottleneck. We present Block Sparse Flash Attention (BSFA), a drop-in replacement that accelerates long-context inference while preserving model quality. Unlike methods that predict importance before computing scores, BSFA computes exact query-key similarities to select the top-k most important value blocks for each query. By comparing per-block maximum scores against calibrated thresholds, we skip approximately 50% of the computation and memory transfers for pruned blocks. Our training-free approach requires only a one-time threshold calibration on a small dataset to learn the per-layer and per-head attention score distributions. We provide a CUDA kernel implementation that can be used as a drop-in replacement for FlashAttention. On Llama-3.1-8B, BSFA achieves up to 1.13x end-to-end speedup on LongBench with only a 1.1% accuracy drop, and up to 1.24x on Needle-in-a-Haystack retrieval at a 1% accuracy drop. The attention kernel itself accelerates by up to 1.38x. We compare BSFA against five recent sparse attention baselines (SpargeAttention, MInference, FlexPrefill, XAttention, and BLASST), and verify the method on Qwen2.5-7B and on A6000 and H100 GPUs. The implementation is available at https://github.com/Danielohayon/Block-Sparse-Flash-Attention.
Daniel Ohayon, Itay Lamprecht, Itay Hubara +3
Technion – Israel Institute of Technology, Haifa, Israel · Intel - Habana Labs · Intel – Habana Labs, Tel Aviv, Israel
FlashAttention improves efficiency through tiling, but its online softmax still relies on floating-point arithmetic for numerical stability, making full quantization difficult. We identify three main obstacles to integer-only FlashAttention: (1) scale explosion during tile-wise accumulation, (2) inefficient shift-based exponential operations on GPUs, and (3) quantization granularity constraints requiring uniform scales for integer comparison. To address these challenges, we propose \textit{QFlash}, an end-to-end integer FlashAttention design that performs softmax entirely in the integer domain and runs as a single Triton kernel. On seven attention workloads from ViT, DeiT, and Swin models, QFlash achieves up to 6.73× speedup over I-ViT and up to 8.69× speedup on Swin, while reducing energy consumption by 18.8% compared to FP16 FlashAttention, without sacrificing Top-1 accuracy on ViT/DeiT and remaining competitive on Swin under per-tensor quantization. Our code is publicly available at https://github.com/EfficientCompLab/qflash.
Sehyeon Oh, Yongin Kwon, Jemin Lee
University of Science and Technology, Daejeon, Republic of Korea · Electronics and Telecommunications Research Institute, Daejeon, Republic of Korea · Pusan National University, Busan, Republic of Korea +1
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose \textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .
Jiatong Ding, Bingxin Xing, Yu Zhang +9
Shanghai Jiao Tong University · Xi’an Jiaotong University · Xiamen University +1