cs.LGSep 30, 2026

LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators

Authors: Stanislav Budzinskiy, Marian Gloser, Tolunay Yilmaz, Ying Hong Tham, Yuanyi Lin, Wenyi Fang, Fan Wu, Philipp Petersen

Organizations: Faculty of Mathematics University of Vienna, Austria · Huawei Heisenberg Research Center, Munich, Germany · Huawei Technologies Co. Ltd

Abstract

While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon. We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention. Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats. We propose the specifications for a dedicated accelerator capable of executing this pipeline efficiently. Simulated experiments with Qwen3 and Gemma 3 show that rerouting a selective minority of sub-blocks to high precision is sufficient to recover the baseline model performance.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Block Sparse Flash Attention

    Dec 7, 2025Daniel Ohayon, Itay Lamprecht, Itay Hubara +3Block Sparse Flash AttentionEfficient Long-Context Inference

  2. QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention

    Apr 28, 2026Sehyeon Oh, Yongin Kwon, Jemin LeeBlock Sparse Flash AttentionQuantized

  3. HyQuant: Hybrid-Precision Quantization for LLM Attention

    Aug 28, 2026Jiatong Ding, Bingxin Xing, Yu Zhang +9Large Language Model QuantizationQantis