The quadratic cost of self-attention limits scalability to long sequences from multidimensional data. We introduce Tucker bottleneck attention (TuBA), which exploits low-rank tensor structure for efficient global token mixing. TuBA projects hidden tensors into compact Tucker cores, performs multi-head self-attention and linear projections on the cores, and writes updates back to the ambient space, enabling subquadratic computation. Its autoregressive extension combines bidirectional interactions within cores with causal attention across cores. On video prediction and global weather forecasting, TuBA achieves favorable accuracy-efficiency trade-offs over standard and efficient attention and task-specific models. Compared to standard self-attention, TuBA reduces error and computation by up to 24.7% and 66.6% for video prediction and 37.1% and 85.1% for autoregressive weather forecasting, with speedups up to 4.27 times. Low-rank Tucker cores and multi-frame generation also outperform full-rank attention and frame-by-frame generation, respectively.
Figures & tables
Figure 1: Overview of Tucker bottleneck attention. TuBA projects a tensor into a compact Tucker core, applies attention on the core, then reconstructs the output.
Method
Accuracy
Efficiency
MSE ↓
MAE ↓
SSIM ↑
PSNR ↑
LPIPS ↓
Params
FLOPs
FPS
FPS ∗
TuBA
113.21
1499.3
0.9664
33.79
0.0339
11.1M
63.5G
580
739
SimVP
116.91
1525.9
0.9824
33.72
0.0340
11.1M
74.0G
589
576
ViT
141.64
1642.6
0.9815
33.50
0.0335
11.1M
172.8G
151
694
Linear
149.79
1639.4
0.9615
32.54
0.0397
11.1M
70.6G
555
741
Linformer
277.30
2107.8
0.9501
30.02
0.0406
23.7M
76.2G
552
701
Table 1: Comparison of predictive accuracy and computational efficiency on Human3.6M.
Method
Accuracy
Efficiency
MSE ↓
MAE ↓
SSIM ↑
PSNR ↑
LPIPS ↓
Params
FLOPs
FPS
FPS ∗
TuBA
156.07
1641.7
0.8963
27.25
0.0772
7.0M
29.2G
142
382
SimVP
156.94
1683.9
0.9019
26.87
0.0676
8.6M
60.6G
218
197
ViT
207.33
1970.4
0.8621
25.85
0.1159
7.0M
87.5G
62
392
Linear
174.18
1732.0
0.8782
26.70
0.1026
7.0M
25.2G
293
417
Linformer
285.74
2303.4
0.8337
24.29
0.1346
16.8M
29.6G
287
381
Table 2: Comparison of predictive accuracy and computational efficiency on KittiCaltech.
Figure 2: Comparison of prediction error (MSE) and computational cost (FLOPs) on Human3.6M and KittiCaltech; the FLOPs axis is compressed above 200G. TuBA lies on the FLOPs–MSE Pareto frontier on both benchmarks.
Figure 3: Effect of Tucker core size on TuBA’s prediction accuracy on KittiCaltech, measured by MSE and SSIM.
Figure 4: Daily NRMSE of autoregressive forecasts over ten days. The dotted line marks the three-day forecast horizon used during training.
Method
NRMSE ↓
z 500 ↓
t 850 ↓
t 2m ↓
u 10 ↓
v 10 ↓
Params
FLOPs
clips/s
TuBA
0.399
0.141
0.162
0.126
0.536
0.627
8.72M
40.8G
600
SimVP
0.373
0.127
0.158
0.122
0.484
0.572
8.60M
67.8G
192
Earthformer
0.415
0.148
0.182
0.194
0.526
0.613
8.66M
86.5G
191
ViT
0.492
0.212
0.204
0.147
0.671
0.793
8.71M
198G
362
Linformer
0.516
0.216
0.212
0.147
0.720
0.866
33.87M
56.7G
434
Linear
0.550
0.259
0.242
0.169
0.731
0.879
8.71M
45.4G
514
Table 3: Comparison of predictive accuracy and computational efficiency on ERA5. Models are trained on 2011–2015 and evaluated on 2017–2018. (AR) denotes autoregressive generation.
Granularity
NRMSE ↓
Params
FLOPs
clips/s
1
0.426
7.27M
89.3G
100
2
0.358
7.41M
85.2G
163
4
0.349
7.67M
67.2G
265
6
0.368
7.93M
62.9G
317
Table 4: Effect of causal granularity (frames per tensor) on autoregressive TuBA’s accuracy and computational efficiency on ERA5.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Tucker reconstruction error of held-out ViT hidden representations on Human3.6M (top) and KittiCaltech (bottom) at three representative layers. Shared factors with per-sample cores closely approach per-sample factors; dotted lines mark TuBA’s core sizes. Random factors are included as a control.
R
Ranks
MSE ↓
MAE ↓
SSIM ↑
PSNR ↑
LPIPS ↓
720
(9,8,10)
170.54
1737.3
0.8889
26.74
0.0881
891
(9,9,11)
165.20
1665.5
0.8881
26.97
0.0898
972
(9,9,12)
172.03
1757.1
0.8855
26.66
0.0972
1080
(9,10,12)
162.53
1641.4
0.8948
27.15
0.0836
1188
(9,11,12)
158.15
1654.2
0.8957
27.14
0.0832
1287
(9,11,13)
156.07
1641.7
0.8963
27.25
0.0772
Appendix
Table 5: Effect of Tucker ranks on TuBA’s accuracy on KittiCaltech.
Human3.6M
KittiCaltech
ERA5
Method
Memory
Time/epoch
Memory
Time/epoch
Memory
Time/epoch
TuBA
10.22
29.3
15.05
1.68
11.99
7.50
ViT
38.72
95.7
52.28
7.23
61.61
12.68
Linear
10.61
23.7
10.86
1.10
13.35
7.67
Linformer
10.47
25.2
10.04
1.30
13.42
7.37
Nyström
10.96
29.5
12.54
1.57
12.96
7.31
Appendix
Table 6: Peak training GPU memory (GiB) and wall-clock time per epoch (minutes) on a single H100 GPU in fp32 . Batch sizes are 4 for Human3.6M and 8 for KittiCaltech and ERA5.
Linear attention and state-space models provide linear-time sequence modeling, but their recurrent memory remains a second-order tensor (a matrix), limiting the order of interactions that can be represented in the state. We introduce the RunningTensor, which generalizes this memory to an order-o tensor, updated by a rank-1 outer product and read by contracting against o−1 vector queries. Order 2 recovers linear attention; we study order 3 as a proof of concept, retaining both recurrent and parallel forms while remaining linear in sequence length T and improving working memory capacity from O(W2) to O(Wo). On synthetic multi-query associative recall, RunningTensor outperforms linear-attention and SSM baselines. After pretraining, it also improves performance on language-understanding and non-synthetic retrieval tasks, suggesting that higher-order recurrent state can provide useful additional memory capacity beyond matrix-valued state.
The O(N2) complexity of attention over N tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces this to O(MN) by representing keys with M codewords, but applies uniform codebook capacity regardless of where attention mass concentrates: high-attention regions of key space may be coarsely approximated while low-attention regions waste representational capacity. We propose Adaptive Vector-Quantized (AVQ) Attention, which adaptively allocates codebook capacity based on attention importance. Starting from a small set of codewords, our method identifies the most important codes during the forward pass and refines them with pre-learned child codewords, achieving fine-grained quantization where it matters most while maintaining coarse quantization elsewhere. We develop an implementation using custom Triton kernels that enables the full adaptive refinement process, including importance scoring, child codeword insertion, and parent contribution replacement, to be carried out within the tiled computation paradigm of Flash Attention with minimal overhead. Our approach maintains O(MN) complexity while achieving improved accuracy-efficiency trade-offs compared to fixed-codebook VQ-attention.
Winfried van den dool, Patrick Forré, Amir Habibian +2
QUVA Lab, University of Amsterdam, The Netherlands · AMLab, Informatics Institute, University of Amsterdam, The Netherlands · AI4Science Lab, University of Amsterdam, The Netherlands +3
Under modern test-time compute and agentic paradigms, language models process ever-longer sequences. Efficient text generation with transformer architectures is increasingly constrained by the Key-Value cache memory footprint and bandwidth. To address this limitation, we introduce Self-Pruned Key-Value Attention (SP-KV), a mechanism designed to predict future KV utility in order to reduce the size of the long-term KV cache. This strategy operates at a fine granularity: a lightweight utility predictor scores each key-value pair, and while recent KVs are always available via a local window, older pairs are written in the cache and used in global attention only if their predicted utility surpasses a given threshold. The LLM and the utility predictor are trained jointly end-to-end exclusively through next-token prediction loss, and are adapted from pretrained LLM checkpoints. Rather than enforcing a fixed compression ratio, SP-KV performs dynamic sparsification: the mechanism adapts to the input and typically reduces the KV cache size by a factor of 3 to 10×, longer sequences often being more compressible. This leads to vast improvements in memory usage and decoding speed, with little to no degradation of validation loss nor performance on a broad set of downstream tasks. Beyond serving as an effective KV-cache reduction mechanism, our method reveals structured layer- and head-specific sparsity patterns that we can use to guide the design of hybrid local-global attention architectures.