Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
Organizations: School of Computer Science and Technology, Soochow University · Baidu Inc, China
Abstract
The quadratic computational complexity of standard attention mechanisms presents a severe scalability bottleneck for LLMs in long-context scenarios. While hybrid attention mechanisms combining Full Attention (FA) and Sparse Attention (SA) offer a potential solution, existing methods typically rely on static allocation ratios that fail to accommodate the variable retrieval demands of different tasks. Furthermore, head-level dynamic sparsity often introduces severe computational load imbalance and synchronization long-tails, which hinder hardware acceleration during autoregressive decoding. To bridge this gap, we introduce Flux Attention, a context-aware framework that dynamically optimizes attention computation at the layer level. By integrating a lightweight Layer Router into frozen pretrained LLMs, the proposed method adaptively routes each layer to FA or SA based on the input context. This layer-wise routing preserves high-fidelity information retrieval while ensuring contiguous memory access, translating theoretical computational reductions into practical wall-clock speedups. As a parameter-efficient approach, our framework requires only 12 hours of training on 8A800 GPUs. Extensive experiments across multiple long-context and mathematical reasoning benchmarks demonstrate that Flux Attention achieves a superior trade-off between performance and inference speed compared with baseline models, with speed improvements of up to and in the prefill and decode stages.
Figures & tables
| Method | S-Doc QA | M-Doc QA | Summ | In-Context | Synthetic | Code | Avg. | ||||||||
| Qasper | MF-en | HotQA | 2Wiki | Gov. | M.News | TREC | TQA | SAMS | PCount | PRe | RB-P | Lcc | Perf. | ||
| Qwen3-4B backbone model | |||||||||||||||
| Qwen3-4B | 35.21 | 52.16 | 44.81 | 32.15 | 33.47 | 23.45 | 70.67 | 88.22 | 39.74 | 2.33 | 96.84 | 50.84 | 57.93 | 48.45 | - |
| + DuoAttention | 35.83 | 49.84 | 47.09 | 32.24 | 33.32 | 23.70 | 69.33 | 85.87 | 39.75 | 4.50 | 94.57 | 50.56 | 57.43 | 48.22 | 0.50 |
| + PruLong | 34.15 | 50.78 | 44.48 | 32.89 | 32.96 | 23.53 | 67.67 | 88.69 | 39.55 | 3.17 | 90.17 | 49.00 | 54.07 | 47.16 | 0.50 |
| + TriangleMix | 35.55 | 52.02 | 45.37 | 31.76 | 33.32 | 23.70 | 69.00 | 88.20 | 39.74 | 3.83 | 91.51 | 48.58 | 56.38 | 47.72 | 0.50 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Selection unit | Selection | Primary phase | Deployment objective |
|---|---|---|---|---|
| Flux Attention (ours) | Whole layer | Learned, input-conditioned FA/SA | Prefill routing, decode reuse | Decode-regular execution, KV memory |
| Elastic Attention [ 42 ] | Head | Learned, test-time adaptive | Prefill/TTFT | Adaptive sparsity ratio |
| DuoAttention [ 48 ] | Head | Offline static allocation | Decode | Hybrid long-context attention |
| PruLong [ 4 ] | Head | Offline static allocation | Decode | Pruned long-context attention |
| TriangleMix [ 15 ] | Block pattern | Fixed triangle topology | Prefill | Training-free sparse backend |
| XAttention [ 53 ] | Block | Fixed block selection score | Prefill | Prefill acceleration |
| Source | Task family it covers | Regime exposed to the router |
|---|---|---|
| ChatQA2-Long-SFT-data [ 52 ] | Single-document and multi-document QA over long inputs | Retrieval-intensive |
| MuSiQue [ 44 ] | Multi-hop question answering | Retrieval-intensive |
| CoLT-132K [ 24 ] | Repository-level code completion | Context-holistic |
| GovReport [ 17 ] | Long-form summarization of government reports | Context-holistic |
| XSum [ 34 ] | Abstractive summarization | Context-holistic |
| Method | Routing / selection unit | Training or calibration in our comparison | Deployment focus |
|---|---|---|---|
| FluxAttn (ours) | Whole layer, input-conditioned FA/SA | Router trained on the shared data mixture; backbone frozen | Decode-regular layer execution |
| PruLong | Offline head selection | Same data and environment; original method hyperparameters | Optimized block-sparse decode |
| DuoAttention | Head allocation | Same data and environment; original method hyperparameters | Hybrid long-context attention |
| TriangleMix / TA | Triangle Attention operator | Training-free | Sparse-attention backend |
| Item | Setting |
|---|---|
| GPU | 1 NVIDIA A800 80GB (PCIe); batched-serving results in Table 14 on 8 NVIDIA H20 |
| Precision | BF16 |
| Batch size | 1 for profiling; 8/16/24/32 for the serving experiment |
| Dense attention path | FlashAttention-2 with the same PyTorch/CUDA environment as the sparse and hybrid paths |
| Sparse attention kernel | Block-Sparse-Attention, block size 64, sink size 128, chunk size 16384 ( Section F.4 ) |
| Warm-up / profiling | 10 warm-up iterations, 50 profiling iterations, one token per step |
| Hyperparameter | Value |
|---|---|
| Model & Training | |
| Base Model | Qwen, Llama |
| Sequence length | 65536 |
| Precision | bfloat16 |
| Global Batch Size | 48 |
| Training Steps | 300 |
| Method | S-Doc QA | M-Doc QA | Summ | In-Context | Synthetic | Code | Avg. | ||||||||
| Qasper | MF-en | HotQA | 2Wiki | Gov. | M.News | TREC | TQA | SAMS | PCount | PRe | RB-P | Lcc | Perf. | ||
| Qwen3-4B backbone model | |||||||||||||||
| Qwen3-4B | 35.21 | 52.16 | 44.81 | 32.15 | 33.47 | 23.45 | 70.67 | 88.22 | 39.74 | 2.33 | 96.84 | 50.84 | 57.93 | 48.45 | - |
| + XAttention | 32.80 | 50.36 | 44.54 | 33.17 | 33.79 | 23.73 | 69.00 | 87.44 | 39.90 | 3.42 | 74.59 | 49.42 | 59.62 | 46.44 | - |
| + LycheeDecode | 34.12 | 52.63 | 44.74 | 31.53 | 31.16 | 23.24 | 68.67 | 87.69 | 39.12 | 4.00 | 96.87 | 49.17 | 57.29 | 47.83 | 0.50 |
| + DuoAttention | 35.83 | 49.84 | 47.09 | 32.24 | 33.32 | 23.70 | 69.33 | 85.87 | 39.75 | 4.50 | 94.57 | 50.56 | 57.43 | 48.22 | 0.50 |
| Parameter | Configuration Details |
|---|---|
| Context Windows | 8k, 16k, 32k, 64k, 128k, 256k |
| Sample Size | 50 samples per task-length pair |
| Evaluation Tasks | Retrieval (NIAH): Single_{1-3} , MultiKey_{1-3} , MultiQuery , MultiValue QA & Extraction: QA_{1,2} , FWE |
| Models | NS-3 | QA-2 | NMK-3 | NS-2 | FWE | QA-1 | NMQ | NMV | NMK-2 | NS-1 | NMK-1 | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Backbone: Qwen3-4B | ||||||||||||
| Qwen3-4B | 95.02 | 35.59 | 37.54 | 80.00 | 84.47 | 40.14 | 66.19 | 65.91 | 47.39 | 100.00 | 73.05 | 66.00 |
| + LycheeDecode | 72.33 | 30.67 | 5.00 | 64.67 | 71.33 | 27.33 | 55.58 | 50.58 | 25.00 | 100.00 | 57.67 | 50.92 |
| + TriangleMix | 95.00 | 34.67 | 34.33 | 82.67 | 83.11 | 35.67 | 60.42 | 60.00 | 43.67 | 99.33 | 72.33 | 63.74 |
| + DuoAttention | 97.08 | 32.12 | 21.45 | 84.75 | 81.38 | 27.17 | 66.67 | 63.41 | 30.66 | 99.28 | 63.50 | 60.67 |
| + PruLong | 96.34 | 33.80 | 22.44 | 81.51 | 84.87 | 22.93 | 64.87 | 57.87 | 27.04 | 100.00 | 71.03 | 60.25 |
| Method | Sparsity | CWE | VT | Mean |
|---|---|---|---|---|
| Backbone (FA) | 0% | 39.90 | 80.60 | 60.25 |
| FluxAttn | 51% dynamic | 43.07 | 82.60 | 62.84 |
| DuoAttention | 50% static | 42.10 | 82.87 | 62.49 |
| Sparsity | XAttention (XA) | Triangle Attention (TA) | ||||||
|---|---|---|---|---|---|---|---|---|
| Doc QA | Multi-Hop | Summ. | Code | Doc QA | Multi-Hop | Summ. | Code | |
| 0.0 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| 0.2 | 102.78 | 96.26 | 99.97 | 98.54 | 101.76 | 99.10 | 94.86 | 95.63 |
| 0.4 | 101.70 | 92.01 | 101.14 | 102.91 | 100.16 | 95.26 | 92.84 | 97.13 |
| 0.6 | 96.87 | 92.37 | 101.04 | 104.54 | 97.98 | 93.21 | 92.32 | 97.01 |
| 0.7 | 95.44 | 90.16 | 95.85 | 104.02 | 93.21 | 81.86 | 90.57 | 101.11 |
| Task | Original | Seed 456 | Mean Std |
|---|---|---|---|
| Qasper | 45.25 | 45.38 | 45.32 0.07 |
| MF-en | 54.42 | 54.42 | 54.42 0.00 |
| HotQA | 54.54 | 54.54 | 54.54 0.00 |
| 2Wiki | 41.34 | 41.34 | 41.34 0.00 |
| Gov. | 34.54 | 34.71 | 34.63 0.09 |
| M.News | 26.16 | 26.22 | 26.19 0.03 |
| Batch size | FluxAttn (tok/s) | FA2 (tok/s) | FluxAttn / FA2 |
| 8 | 194 | 147 | 1.32 |
| 16 | 196 | 95 | 2.07 |
| 24 | 280 | OOM | — |
| 32 | OOM | OOM | — |
| Metric | 8K | 16K | 32K | 64K | 128K | 256K |
|---|---|---|---|---|---|---|
| Acc. Dense / full | 92.88 | 92.83 | 89.46 | 70.79 | 80.12 | 72.34 |
| Acc. PruLong / head | 86.96 | 76.55 | 70.65 | 54.52 | 48.18 | 30.00 |
| Acc. FluxAttn / layer | 90.11 | 79.39 | 79.22 | 56.08 | 62.94 | 59.39 |
| (PruLong / FluxAttn) | .50 / .53 | .50 / .50 | .50 / .53 | .50 / .53 | .50 / .50 | .50 / .50 |
| Prefill speedup, PruLong | 1.06 | 1.14 | 1.25 | 1.42 | 1.60 | 1.76 |
| Prefill speedup, FluxAttn | 1.22 | 1.18 | 1.36 | 1.59 | 1.67 | 1.77 |
| Ctx | Method / granularity | Acc. | Prefill ms (speedup) | Decode ms/tok (speedup) | Persistent KV GiB (saving) | Peak GiB | |
|---|---|---|---|---|---|---|---|
| 8K | Dense / full | 92.88 | 0 | 985.0 (1.00 ) | 25.17 (1.00 ) | 0.89 (0.0%) | 16.95 |
| PruLong / head | 86.96 | 0.500 | 926.2 (1.06 ) | 21.88 (1.15 ) | 0.89 (0.0%) | 17.03 | |
| FluxAtt / layer | 90.11 | 0.531 | 807.4 (1.22 ) | 21.15 (1.19 ) | 0.75 (16.1%) | 16.89 | |
| 16K | Dense / full | 92.83 | 0 | 2280.1 (1.00 ) | 16.31 (1.00 ) | 1.81 (0.0%) | 18.96 |
| PruLong / head | 76.55 | 0.500 | 2004.3 (1.14 ) | 13.26 (1.23 ) | 1.81 (0.0%) | 19.09 | |
| FluxAtt / layer | 79.39 | 0.500 | 1932.3 (1.18 ) | 12.55 (1.30 ) | 1.28 (28.9%) | 18.29 |
| Method | Granularity | Sparsity | Performance | Speedup | |
|---|---|---|---|---|---|
| LongBench | RULER | ||||
| Elastic Attention | Head-wise | 68% | 53.35 | 72.85 | 2.6 |
| Flux Attention (Ours) | Layer-wise | 51% | 52.28 | 76.75 | 2.5 |
| Model | Method | Performance | |
|---|---|---|---|
| LongBench | RULER | ||
| Qwen3-4B | Dense (Full Attention) | 48.45 | 66.00 |
| Elastic Attention | 48.08 | 61.81 | |
| Flux Attention (Ours) | 48.72 | 66.95 | |
| Qwen3-8B | Dense (Full Attention) | 52.16 | 75.74 |
| Elastic Attention | 51.51 | 71.74 | |
| Task | Spearman | Mean FA |
|---|---|---|
| Qasper | 35.2% | |
| MultiFieldQA | 44.2% | |
| HotpotQA | 38.4% | |
| 2WikiMQA | 39.6% | |
| GovReport | 57.5% | |
| MultiNews | 57.7% |