The quadratic computational complexity of standard attention mechanisms presents a severe scalability bottleneck for LLMs in long-context scenarios. While hybrid attention mechanisms combining Full Attention (FA) and Sparse Attention (SA) offer a potential solution, existing methods typically rely on static allocation ratios that fail to accommodate the variable retrieval demands of different tasks. Furthermore, head-level dynamic sparsity often introduces severe computational load imbalance and synchronization long-tails, which hinder hardware acceleration during autoregressive decoding. To bridge this gap, we introduce Flux Attention, a context-aware framework that dynamically optimizes attention computation at the layer level. By integrating a lightweight Layer Router into frozen pretrained LLMs, the proposed method adaptively routes each layer to FA or SA based on the input context. This layer-wise routing preserves high-fidelity information retrieval while ensuring contiguous memory access, translating theoretical computational reductions into practical wall-clock speedups. As a parameter-efficient approach, our framework requires only 12 hours of training on 8×A800 GPUs. Extensive experiments across multiple long-context and mathematical reasoning benchmarks demonstrate that Flux Attention achieves a superior trade-off between performance and inference speed compared with baseline models, with speed improvements of up to 2.8× and 2.0× in the prefill and decode stages.
Figures & tables
Figure 1 : Impact of sparsity on performance and decoding efficiency. (a) Certain tasks suffer performance collapse beyond a specific threshold; this sweep uses Streaming Sparse Attention (SSA) with the configuration of Appendix F.4 . (b) Layer-level sparsity achieves substantial decoding speedup, while head-level sparsity yields marginal speedup.
Figure 2 : Overview of our dynamic layer-level routing architecture. The model incorporates a Layer Router that assigns each layer to either FA or SA based on the input query xQ .
Method
S-Doc QA
M-Doc QA
Summ
In-Context
Synthetic
Code
Avg.
Qasper
MF-en
HotQA
2Wiki
Gov.
M.News
TREC
TQA
SAMS
PCount
PRe
RB-P
Lcc
Perf.
ΩMSR
Qwen3-4B backbone model
Qwen3-4B
35.21
52.16
44.81
32.15
33.47
23.45
70.67
88.22
39.74
2.33
96.84
50.84
57.93
48.45
-
+ DuoAttention
35.83
49.84
47.09
32.24
33.32
23.70
69.33
85.87
39.75
4.50
94.57
50.56
57.43
48.22
0.50
+ PruLong
34.15
50.78
44.48
32.89
32.96
23.53
67.67
88.69
39.55
3.17
90.17
49.00
54.07
47.16
0.50
+ TriangleMix
35.55
52.02
45.37
31.76
33.32
23.70
69.00
88.20
39.74
3.83
91.51
48.58
56.38
47.72
0.50
Table 1 : Performance on LongBench-E [ 1 ] . We report average performance (Perf.) and ΩMSR per task category. The 1st and the 2nd performance in each comparison group are highlighted with bold font and underlined , respectively. Gray-shaded rows denote the sparse-decode configuration.
Figure 3 : Speedup comparison across different context lengths. The dotted line represents the dense baseline performance (1.0x).
Figure 4 : Overview of the layer-wise routing activation frequencies in Llama-3.1-8B-Instruct. Dark blue indicates layers consistently routed to FA across all six tasks in LongBench-E, whereas light blue denotes layers consistently routed to SA.
Figure 5 : Comparison of performance and test-time ΩMSR among different training sparsity target t settings. The bar chart denotes the performance and the line chart denotes ΩMSR in each task.
Figure 6 : Performance trajectories during continued training with a frozen Layer Router. The backbone effectively adapts its representations to the established sparse pathways, demonstrating steady improvement over time.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Selection unit
Selection
Primary phase
Deployment objective
Flux Attention (ours)
Whole layer
Learned, input-conditioned FA/SA
Prefill routing, decode reuse
Decode-regular execution, KV memory
Elastic Attention [ 42 ]
Head
Learned, test-time adaptive
Prefill/TTFT
Adaptive sparsity ratio
DuoAttention [ 48 ]
Head
Offline static allocation
Decode
Hybrid long-context attention
PruLong [ 4 ]
Head
Offline static allocation
Decode
Pruned long-context attention
TriangleMix [ 15 ]
Block pattern
Fixed triangle topology
Prefill
Training-free sparse backend
XAttention [ 53 ]
Block
Fixed block selection score
Prefill
Prefill acceleration
Appendix
Table 3: Positioning of Flux Attention against representative dynamic and fine-grained attention-selection methods. The axes are the unit at which computation is selected, whether the selection is learned or fixed, the inference phase the method primarily targets, and the deployment objective. Flux Attention is complementary to fine-grained operators: it chooses the coarse per-layer execution mode, while a method such as vAttention can supply fine-grained sparsity inside the routed sparse branch.
Source
Task family it covers
Regime exposed to the router
ChatQA2-Long-SFT-data [ 52 ]
Single-document and multi-document QA over long inputs
Retrieval-intensive
MuSiQue [ 44 ]
Multi-hop question answering
Retrieval-intensive
CoLT-132K [ 24 ]
Repository-level code completion
Context-holistic
GovReport [ 17 ]
Long-form summarization of government reports
Context-holistic
XSum [ 34 ]
Abstractive summarization
Context-holistic
Appendix
Table 4: Training mixture of the Layer Router (approximately 0.74B tokens over 1K–64K sequences). The mixture is chosen to expose the router to both retrieval-intensive and context-holistic regimes; it does not import the evaluation benchmark suites.
Method
Routing / selection unit
Training or calibration in our comparison
Deployment focus
FluxAttn (ours)
Whole layer, input-conditioned FA/SA
Router trained on the shared data mixture; backbone frozen
Decode-regular layer execution
PruLong
Offline head selection
Same data and environment; original method hyperparameters
Optimized block-sparse decode
DuoAttention
Head allocation
Same data and environment; original method hyperparameters
Hybrid long-context attention
TriangleMix / TA
Triangle Attention operator
Training-free
Sparse-attention backend
Appendix
Table 5: Baseline comparison controls. All trainable baselines are trained in the same environment and on the same data mixture as Flux Attention while retaining their original method-specific hyperparameters; TriangleMix’s Triangle Attention (TA) is training-free and serves as a sparse-attention backend for Flux rather than as a routed competitor.
Item
Setting
GPU
1 × NVIDIA A800 80GB (PCIe); batched-serving results in Table 14 on 8 × NVIDIA H20
Precision
BF16
Batch size
1 for profiling; 8/16/24/32 for the serving experiment
Dense attention path
FlashAttention-2 with the same PyTorch/CUDA environment as the sparse and hybrid paths
10 warm-up iterations, 50 profiling iterations, one token per step
Appendix
Table 6: Systems setup for the latency and memory measurements reported in Figure 3 and Table 15 . All compared methods share the hardware, precision, input/output shapes, warm-up, repetition, and timing protocol; the attention execution path is the controlled difference.
Hyperparameter
Value
Model & Training
Base Model
Qwen, Llama
Sequence length
65536
Precision
bfloat16
Global Batch Size
48
Training Steps
300
Appendix
Table 7: Hyperparameters: General configuration.
Method
S-Doc QA
M-Doc QA
Summ
In-Context
Synthetic
Code
Avg.
Qasper
MF-en
HotQA
2Wiki
Gov.
M.News
TREC
TQA
SAMS
PCount
PRe
RB-P
Lcc
Perf.
ΩMSR
Qwen3-4B backbone model
Qwen3-4B
35.21
52.16
44.81
32.15
33.47
23.45
70.67
88.22
39.74
2.33
96.84
50.84
57.93
48.45
-
+ XAttention
32.80
50.36
44.54
33.17
33.79
23.73
69.00
87.44
39.90
3.42
74.59
49.42
59.62
46.44
-
+ LycheeDecode
34.12
52.63
44.74
31.53
31.16
23.24
68.67
87.69
39.12
4.00
96.87
49.17
57.29
47.83
0.50
+ DuoAttention
35.83
49.84
47.09
32.24
33.32
23.70
69.33
85.87
39.75
4.50
94.57
50.56
57.43
48.22
0.50
Appendix
Table 8: Performance on LongBench-E [ 1 ] . We report average performance (Perf.) and ΩMSR per task category. The 1st and the 2nd performance in each comparison group are highlighted with bold font and underlined , respectively. Gray-shaded rows denote the sparse-decode configuration.
Table 9: Detailed configuration for the RULER benchmark evaluation. We evaluate across exponentially increasing context windows up to 256k tokens.
Models
NS-3
QA-2
NMK-3
NS-2
FWE
QA-1
NMQ
NMV
NMK-2
NS-1
NMK-1
Avg.
Backbone: Qwen3-4B
Qwen3-4B
95.02
35.59
37.54
80.00
84.47
40.14
66.19
65.91
47.39
100.00
73.05
66.00
+ LycheeDecode
72.33
30.67
5.00
64.67
71.33
27.33
55.58
50.58
25.00
100.00
57.67
50.92
+ TriangleMix
95.00
34.67
34.33
82.67
83.11
35.67
60.42
60.00
43.67
99.33
72.33
63.74
+ DuoAttention
97.08
32.12
21.45
84.75
81.38
27.17
66.67
63.41
30.66
99.28
63.50
60.67
+ PruLong
96.34
33.80
22.44
81.51
84.87
22.93
64.87
57.87
27.04
100.00
71.03
60.25
Appendix
Table 10: Model performance on RULER tasks. Bold and underline indicate the best and second-best performance within each backbone group, respectively.NS denotes Single-needle in a haystack (NIAH) tasks, NMK stands for Multi-key NIAH, NMV is Multi-value NIAH, and NMQ refers to Multi-query NIAH. QA represents standard Question Answering tasks, and FWE stands for Frequent Words Extraction. Numeric suffixes (e.g., -1, -2, -3) indicate different subsets or difficulty levels within the corresponding task categories.
Method
Sparsity
CWE
VT
Mean
Backbone (FA)
0%
39.90
80.60
60.25
FluxAttn
≈ 51% dynamic
43.07
82.60
62.84
DuoAttention
50% static
42.10
82.87
62.49
Appendix
Table 11: Common Words Extraction (CWE) and Variable Tracking (VT), the two RULER tasks omitted from Table 10 , evaluated on Llama-3.1-8B-Instruct under the same six context lengths {8K,16K,32K,64K,128K,256K} . Including them preserves the main conclusion: FluxAttn remains ahead of both the dense backbone and the static head-level baseline at approximately matched sparsity. This added evaluation covers Llama-3.1-8B-Instruct only.
Sparsity
XAttention (XA)
Triangle Attention (TA)
Doc QA
Multi-Hop
Summ.
Code
Doc QA
Multi-Hop
Summ.
Code
0.0
100.00
100.00
100.00
100.00
100.00
100.00
100.00
100.00
0.2
102.78
96.26
99.97
98.54
101.76
99.10
94.86
95.63
0.4
101.70
92.01
101.14
102.91
100.16
95.26
92.84
97.13
0.6
96.87
92.37
101.04
104.54
97.98
93.21
92.32
97.01
0.7
95.44
90.16
95.85
104.02
93.21
81.86
90.57
101.11
Appendix
Table 12: Sensitivity of the Figure 1 (a) observation to the sparse operator, on Llama-3.1-8B-Instruct. Values are performance normalized to the dense configuration (100.00) and are averaged over the LongBench sub-tasks grouped below. Both XAttention (XA) and Triangle Attention (TA) reproduce the qualitative trend of Figure 1 (a)—retrieval-heavy QA degrades sharply at high sparsity while summarization and code remain tolerant—but the onset and magnitude of the degradation are backend-dependent, which is why the router conditions its allocation on the chosen sparse branch instead of assuming a universal threshold.
Task
Original
Seed 456
Mean ± Std
Qasper
45.25
45.38
45.32 ± 0.07
MF-en
54.42
54.42
54.42 ± 0.00
HotQA
54.54
54.54
54.54 ± 0.00
2Wiki
41.34
41.34
41.34 ± 0.00
Gov.
34.54
34.71
34.63 ± 0.09
M.News
26.16
26.22
26.19 ± 0.03
Appendix
Table 13: LongBench-E per-task results for two independently trained FA-SSA routers on Llama-3.1-8B-Instruct. Router-training stochasticity shifts the task averages by at most 0.5 points, and the aggregate difference between the two runs is 0.08 points.
Batch size
FluxAttn (tok/s)
FA2 (tok/s)
FluxAttn / FA2
8
194
147
1.32 ×
16
196
95
2.07 ×
24
280
OOM
—
32
OOM
OOM
—
Appendix
Table 14: Decode throughput under batched serving on Llama-3.1-8B-Instruct at a 32K context, measured on a node of 8 × NVIDIA H20 GPUs. Throughput is aggregate tokens/s across the served batch. OOM means the configuration exhausts GPU memory on this hardware; we report it rather than extrapolating throughput. The relative advantage does not disappear once heterogeneous requests are co-scheduled, and FluxAttn reaches a larger feasible batch before running out of memory.
Metric
8K
16K
32K
64K
128K
256K
Acc. Dense / full
92.88
92.83
89.46
70.79
80.12
72.34
Acc. PruLong / head
86.96
76.55
70.65
54.52
48.18
30.00
Acc. FluxAttn / layer
90.11
79.39
79.22
56.08
62.94
59.39
ΩMSR (PruLong / FluxAttn)
.50 / .53
.50 / .50
.50 / .53
.50 / .53
.50 / .50
.50 / .50
Prefill speedup, PruLong
1.06 ×
1.14 ×
1.25 ×
1.42 ×
1.60 ×
1.76 ×
Prefill speedup, FluxAttn
1.22 ×
1.18 ×
1.36 ×
1.59 ×
1.67 ×
1.77 ×
Appendix
Table 15: Quality–latency–memory trade-off at matched sparsity on Llama-3.1-8B-Instruct. Accuracy is the RULER score per context length; speedups are relative to the dense FA2 baseline at the same context length, with prefill and decode reported separately because they are different phases. FluxAttn and PruLong operate at approximately matched model sparsity ( ΩMSR≈0.50 ), so the comparison isolates routing granularity rather than the amount of computation retained. “KV saving” is the reduction in persistent decode-time KV memory relative to dense. The accuracy entries coincide with the corresponding rows of Table 2 , where the FluxAttn configuration is the sparse-decode one. Raw latencies and peak memory are given in Table 16 , and the systems setup in Table 6 .
Ctx
Method / granularity
Acc.
ΩMSR
Prefill ms (speedup)
Decode ms/tok (speedup)
Persistent KV GiB (saving)
Peak GiB
8K
Dense / full
92.88
0
985.0 (1.00 × )
25.17 (1.00 × )
0.89 (0.0%)
16.95
PruLong / head
86.96
0.500
926.2 (1.06 × )
21.88 (1.15 × )
0.89 (0.0%)
17.03
FluxAtt / layer
90.11
0.531
807.4 (1.22 × )
21.15 (1.19 × )
0.75 (16.1%)
16.89
16K
Dense / full
92.83
0
2280.1 (1.00 × )
16.31 (1.00 × )
1.81 (0.0%)
18.96
PruLong / head
76.55
0.500
2004.3 (1.14 × )
13.26 (1.23 × )
1.81 (0.0%)
19.09
FluxAtt / layer
79.39
0.500
1932.3 (1.18 × )
12.55 (1.30 × )
1.28 (28.9%)
18.29
Appendix
Table 16: Full quality–latency–memory Pareto comparison at matched granularity controls on Llama-3.1-8B-Instruct; the compact view is Table 15 . All methods are measured under the same precision and profiling procedure (Appendix F.5 ); speedups are relative to the dense FA2 baseline at the same context length. “Persistent KV” is the retained decode-time KV state and “Peak” the peak GPU memory. At 128K and 256K, where ΩMSR is exactly matched at 0.50 , FluxAtt improves accuracy over PruLong by 14.76 and 29.39 points, respectively, while reducing decode latency by 19.2% and 32.0% relative to PruLong. Prefill and decode columns report different inference phases and must not be conflated: head-level sparsity is competitive on prefill while remaining behind on per-token decode.
Method
Granularity
Sparsity
Performance
Speedup
LongBench
RULER
Elastic Attention
Head-wise
68%
53.35
72.85
2.6 ×
Flux Attention (Ours)
Layer-wise
51%
52.28
76.75
2.5 ×
Appendix
Table 17: Performance and efficiency comparison between Elastic Attention and Flux Attention on Llama-3.1-8B-Instruct. Flux Attention operates at a substantially lower sparsity level (51% vs. 68%), i.e., it retains more attention computation, and still achieves a comparable prefill/TTFT speedup (2.5 × vs. 2.6 × ) thanks to its hardware-friendly layer-wise routing. The reported speedups are prefill speedups; the decode-phase comparison is given in Table 15 . Retaining more attention computation allows Flux Attention to deliver superior performance on the rigorous RULER benchmark while maintaining competitive results on LongBench.
Model
Method
Performance
LongBench
RULER
Qwen3-4B
Dense (Full Attention)
48.45
66.00
Elastic Attention
48.08
61.81
Flux Attention (Ours)
48.72
66.95
Qwen3-8B
Dense (Full Attention)
52.16
75.74
Elastic Attention
51.51
71.74
Appendix
Table 18: Comprehensive performance comparison among Dense Attention, Elastic Attention, and our Flux Attention across various foundation models. Flux Attention effectively mitigates the severe performance degradation seen in Elastic Attention on the RULER benchmark while maintaining highly competitive results on LongBench.
Figure 7: Evolution of sparsity levels across training steps under different data distributions. Left: Training on a well-balanced dataset, where the router successfully disentangles tasks into distinct sparsity levels. Right: Training on an unbalanced dataset dominated by context-holistic tasks, leading to homogenized routing.
Figure 8: Impact of pooling window size on downstream performance and routing sparsity ( ΩMSR ). We evaluate varying truncation budgets ( L∈{50,100,200,400,800,Full} ), retaining only the sequence boundaries (prefix and suffix). Increasing the pooling size beyond 100 tokens introduces context noise, which disrupts the routing mechanism. Consequently, the router misclassifies task features and assigns excessive sparsity to retrieval-intensive tasks, thereby degrading the overall performance.
Task
Spearman ρ
Mean FA
Qasper
+0.027
35.2%
MultiFieldQA
−0.250
44.2%
HotpotQA
−0.205
38.4%
2WikiMQA
−0.167
39.6%
GovReport
−0.293
57.5%
MultiNews
−0.314
57.7%
Appendix
Table 19: Corrected correlation between the task-independent offline entropy score Eℓ (Appendix D.1 ) and the learned soft FA probability of each layer, on Llama-3.1-8B-Instruct. Eℓ is computed once from a trace-normalized hidden-state Gram matrix using its top-128 eigenvalues on a profiling split of 26 disjoint task-balanced LongBench-E samples; routing is collected over 20 samples from each of 13 LongBench-E tasks (260 samples total) across 32 layers. Correlations are weak to moderate and task-dependent, so the learned router departs from a static entropy ranking rather than reproducing it; entropy is not used as a routing label or supervision target. The macro-average Pearson correlation is r=−0.108 , and the macro-average Spearman correlation is stable to the eigenvalue-truncation threshold ( −0.231 , −0.234 , −0.284 , −0.215 for K=16,32,64,128 ).
Figure 9: Router latency analysis. The router incurs negligible overhead (avg. 0.20 ms). Our design ensures length-invariant stability, maintaining constant speed from 512 to 1M tokens.
Figure 10: Decomposition of Training Objectives for Flux Attention. We visualize the training dynamics of the Layer Router, separating the total loss into (a) the primary language modeling objective and (b) the sparsity regularization term. Subfigures (c) and (d) illustrate the task-level differentiation in sparsity allocation ( ΩMSR ) and adaptive coefficients ( λ ), demonstrating how the model automatically distinguishes between context-holistic and retrieval-intensive tasks.
Figure 11: Comparison on a long-context reading comprehension task. Our model accurately extracts and verifies the severity statistics of outdated cooking methods in Africa compared to global figures, while all baselines consistently fall for the same unsupported distractor regarding carbon markets.
Figure 12: Qualitative comparison on identifying the core argument in a philosophical legal text. Our model successfully synthesizes the text to identify the underlying argumentative strategy (refutation via analogy), whereas baselines are easily distracted by literal sentences from the title and opening hook.
Figure 13: Qualitative comparison on extracting technical methodology from a machine learning paper. Our model accurately identifies the specific bounding box encoding strategy, whereas all baselines suffer from hallucination, confidently generating plausible but incorrect architectural details (Fourier embeddings) not supported by the text.