Sequential Functional Structured Tucker Compression for Large Language Model Attentions
Organizations: University of Manitoba · McGill University · Simpleway · The University of Hong Kong · McMaster University
Abstract
Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.
Figures & tables
| Model | Ftc | Seq-SVD | enh-L | TLLM | MPIFA | SVD | Basis | Ftc | Seq-SVD | enh-L | TLLM | MPIFA | SVD | Basis |
| GPT-J-6B (8.86) | 9.30 | 9.47 | 9.93 | 9.68 | 9.46 | 10.26 | 10.16 | 13.60 | 14.84 | 17.47 | 19.22 | 17.24 | 25.04 | 23.02 |
| Llama-2-13B (4.89) | 5.11 | 5.27 | 5.13 | 5.52 | 5.26 | 5.68 | 5.57 | 7.52 | 9.00 | 9.79 | 13.01 | 10.21 | 14.20 | 12.77 |
| Mistral-7B (5.32) | 5.44 | 5.69 | 5.58 | 6.07 | 5.63 | 5.92 | 5.75 | 7.33 | 12.10 | 146.31 | 18.41 | 15.92 | 25.40 | 16.18 |
| Llama-3-8B (6.14) | 6.64 | 7.03 | 6.82 | 8.70 | 7.26 | 8.02 | 7.60 | 10.44 | 19.46 | 20.84 | 29.85 | 27.02 | 48.09 | 35.60 |
| Qwen2.5-7B (6.85) | 7.20 | 7.98 | 7.76 | 34.10 | 7.84 | 8.56 | 7.82 | 12.55 | 150.93 | 106.59 | 188.22 | 7791 | 255.23 | 140.65 |
| Ftc | Seq-SVD | enh-LeSTD | TLLM-gen | MPIFA | SVD-LLM | Basis | |||||||||
| Model | a6 | G8K | a6 | G8K | a6 | G8K | a6 | G8K | a6 | G8K | a6 | G8K | a6 | G8K | |
| Mistral-7B (0.740 / 0.285) | 0.6 | 0.726 | 0.24 | 0.684 | 0.17 | 0.707 | 0.15 | 0.654 | 0.10 | 0.695 | 0.20 | 0.667 | 0.17 | 0.688 | 0.20 |
| 0.4 | 0.689 | 0.10 | 0.569 | 0.01 | 0.625 | 0.01 | 0.501 | 0.01 | 0.552 | 0.02 | 0.505 | 0.02 | 0.549 | 0.03 | |
| 0.2 | 0.512 | 0.01 | 0.359 | 0.00 | 0.320 | 0.00 | 0.342 | 0.01 | 0.314 | 0.01 | 0.307 | 0.01 | 0.294 | 0.00 | |
| Llama-3-8B (0.735 / 0.435) | 0.6 | 0.716 | 0.29 | 0.642 | 0.09 | 0.694 | 0.20 | 0.597 | 0.04 | 0.639 | 0.01 | 0.575 | 0.02 | 0.615 | 0.04 |
| 0.4 | 0.668 | 0.10 | 0.510 | 0.04 | 0.601 | 0.04 | 0.447 | 0.03 | 0.448 | 0.02 | 0.410 | 0.04 | 0.474 | 0.03 | |
| Model | full | w/o whitening | w/o slice-wise norm | w/o seq. calibration | w/o anchoring | K/V repeat-expand | dense core | no joint tensor | w/o target | |
| Mistral-7B | 0.6 | 5.440 | 5.56 (+2.2%) | 5.50 (+1.1%) | 5.46 (+0.4%) | 5.44 (+0.0%) | 5.53 (+1.6%) | 6.37 (+17.2%) | 5.69 (+4.7%) | 5.46 (+0.4%) |
| Mistral-7B | 0.2 | 7.334 | 13.19 (+79.8%) | 9.33 (+27.2%) | 13.24 (+80.6%) | 7.45 (+1.5%) | 10.41 (+41.9%) | 7.65 (+4.3%) | 12.10 (+65.0%) | 10.15 (+38.4%) |
| Llama-3-8B | 0.6 | 6.640 | 6.73 (+1.4%) | 7.01 (+5.5%) | 6.94 (+4.5%) | 6.60 (-0.6%) | 6.87 (+3.4%) | 8.45 (+27.3%) | 7.03 (+5.9%) | 6.94 (+4.5%) |
| Llama-3-8B | 0.2 | 10.440 | 20.56 (+96.9%) | 19.77 (+89.4%) | 33.65 (+222.4%) | 10.87 (+4.1%) | 17.18 (+64.6%) | 11.19 (+7.2%) | 19.46 (+86.4%) | 29.50 (+182.6%) |
| Qwen2.5-7B | 0.6 | 7.197 | 7.77 (+7.9%) | 7.17 (-0.3%) | 7.24 (+0.5%) | 7.24 (+0.5%) | 7.24 (+0.6%) | 12.71 (+76.6%) | 7.98 (+10.9%) | 7.19 (-0.1%) |
| Qwen2.5-7B | 0.2 | 12.548 | 221.15 (+1662.4%) | 74.34 (+492.4%) | 21.36 (+70.2%) | 14.36 (+14.5%) | 12.88 (+2.6%) | 26.89 (+114.3%) | 150.93 (+1102.8%) | 12.88 (+2.6%) |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Whole-model storage removed (%) | |||||
| Model | Attn. share | ||||
| GPT-J-6B | 31.1% | 6.2 | 12.4 | 18.7 | 24.9 |
| Llama-2-13B | 32.2% | 6.4 | 12.9 | 19.3 | 25.8 |
| Mistral-7B | 18.5% | 3.7 | 7.4 | 11.1 | 14.8 |
| Llama-3-8B | 16.7% | 3.3 | 6.7 | 10.0 | 13.4 |
| Qwen2.5-7B | 10.8% | 2.2 | 4.3 | 6.5 | 8.6 |
| Model | mode | 0.8 | 0.6 | 0.4 | 0.2 |
| GPT-J-6B | 1 card | 38 | 34 | 26 | 21 |
| Llama-2-13B | 1 card | 166 | 93 | 75 | 39 |
| Mistral-7B | 1 card | 33 | 30 | 26 | 26 |
| Llama-3-8B | 1 card | 34 | 29 | 27 | 25 |
| Qwen2.5-7B | 1 card | 21 | 20 | 16 | 16 |
| Qwen3-8B | 1 card | 39 | 33 | 28 | 27 |
| Model | Ftc | Seq-SVD | enh-LeSTD | TLLM-gen | MPIFA | SVD-LLM | Basis | |
| GPT-J-6B | 0.8 | 9.04 | 9.14 | 8.99 | 9.24 | 9.04 | 9.60 | 9.61 |
| 0.6 | 9.30 | 9.47 | 9.93 | 9.68 | 9.46 | 10.26 | 10.16 | |
| 0.4 | 10.03 | 10.46 | 10.31 | 11.05 | 10.68 | 12.16 | 11.80 | |
| 0.2 | 13.60 | 14.84 | 17.47 | 19.22 | 17.24 | 25.04 | 23.02 | |
| Llama-2-13B | 0.8 | 4.98 | 5.06 | 4.92 | 5.13 | 5.00 | 5.31 | 5.28 |
| 0.6 | 5.11 | 5.27 | 5.13 | 5.52 | 5.26 | 5.68 | 5.57 |
| PTB | C4 | ||||||||||||||
| Model | F | qS | eL | TG | MP | SL | BS | F | qS | eL | TG | MP | SL | BS | |
| GPT-J-6B | 0.8 | 16.33 | 16.80 | 16.28 | 17.03 | 16.30 | 17.45 | 17.42 | 13.62 | 14.12 | 13.29 | 14.06 | 13.62 | 14.84 | 14.83 |
| 0.6 | 17.20 | 18.39 | 17.85 | 18.43 | 17.46 | 19.74 | 19.44 | 14.33 | 15.24 | 14.18 | 15.09 | 14.82 | 16.45 | 16.27 | |
| 0.4 | 20.34 | 25.08 | 21.93 | 25.25 | 24.20 | 30.19 | 28.64 | 16.61 | 18.80 | 16.94 | 18.75 | 18.69 | 21.89 | 21.26 | |
| 0.2 | 50.52 | 67.11 | 82.87 | 101.12 | 88.92 | 162.55 | 115.96 | 29.12 | 34.64 | 40.79 | 42.92 | 42.27 | 65.53 | 59.42 | |
| Llama-2-13B | 0.8 | 52.27 | 61.58 | 53.14 | 54.14 | 56.42 | 61.75 | 60.82 | 7.05 | 7.60 | 7.01 | 7.45 | 7.04 | 8.22 | 8.14 |
| Model | Method | PIQA | ARC-e | ARC-c | HSwag | WinoG | LAMB | MathQA | GSM8K | avg6 | |
| GPT-J-6B | dense | 1.0 | 0.756 | 0.616 | 0.367 | 0.654 | 0.654 | 0.678 | 0.281 | 0.015 | 0.621 |
| Ftc | 0.8 | 0.747 | 0.607 | 0.358 | 0.642 | 0.659 | 0.665 | 0.269 | 0.010 | 0.613 | |
| Ftc | 0.6 | 0.742 | 0.589 | 0.333 | 0.626 | 0.640 | 0.658 | 0.264 | 0.005 | 0.598 | |
| Ftc | 0.4 | 0.735 | 0.577 | 0.317 | 0.567 | 0.618 | 0.573 | 0.263 | 0.025 | 0.564 | |
| Ftc | 0.2 | 0.684 | 0.464 | 0.270 | 0.438 | 0.575 | 0.255 | 0.258 | 0.000 | 0.448 | |
| enh-LeSTD | 0.8 | 0.758 | 0.620 | 0.362 | 0.649 | 0.651 | 0.682 | 0.270 | 0.010 | 0.620 |
| Model | Method | PIQA | ARC-e | ARC-c | HSwag | WinoG | LAMB | MathQA | GSM8K | avg6 | |
| Mistral-7B | dense | 1.0 | 0.814 | 0.807 | 0.530 | 0.796 | 0.740 | 0.753 | 0.367 | 0.285 | 0.740 |
| Ftc | 0.8 | 0.805 | 0.812 | 0.521 | 0.779 | 0.745 | 0.737 | 0.345 | 0.235 | 0.733 | |
| Ftc | 0.6 | 0.805 | 0.811 | 0.506 | 0.765 | 0.741 | 0.728 | 0.336 | 0.240 | 0.726 | |
| Ftc | 0.4 | 0.795 | 0.781 | 0.474 | 0.726 | 0.701 | 0.655 | 0.303 | 0.105 | 0.689 | |
| Ftc | 0.2 | 0.720 | 0.595 | 0.321 | 0.555 | 0.619 | 0.263 | 0.247 | 0.015 | 0.512 | |
| Seq-SVD | 0.6 | 0.781 | 0.759 | 0.450 | 0.711 | 0.716 | 0.690 | 0.320 | 0.175 | 0.684 |
| Model | Method | PIQA | ARC-e | ARC-c | HSwag | WinoG | LAMB | MathQA | GSM8K | avg6 | |
| Qwen3-8B | dense | 1.0 | 0.786 | 0.801 | 0.561 | 0.774 | 0.724 | 0.704 | 0.562 | 0.855 | 0.725 |
| Ftc | 0.8 | 0.788 | 0.832 | 0.583 | 0.763 | 0.735 | 0.664 | 0.563 | 0.765 | 0.728 | |
| Ftc | 0.6 | 0.781 | 0.793 | 0.564 | 0.750 | 0.715 | 0.640 | 0.515 | 0.735 | 0.707 | |
| Ftc | 0.4 | 0.730 | 0.641 | 0.441 | 0.712 | 0.670 | 0.553 | 0.316 | 0.105 | 0.624 | |
| Ftc | 0.2 | 0.690 | 0.550 | 0.324 | 0.550 | 0.582 | 0.305 | 0.247 | 0.005 | 0.500 | |
| Seq-SVD | 0.6 | 0.765 | 0.712 | 0.487 | 0.718 | 0.699 | 0.627 | 0.423 | 0.455 | 0.668 |
| , wt2 ppl | no QKV map (final) | ||||||
| Mistral-7B | 8.10 | 8.13 | 8.04 | 7.85 | 7.64 | 7.91 | 7.33 |
| Llama-3-8B | 12.40 | 12.35 | 11.87 | 11.32 | 11.18 | 11.37 | 10.44 |
| Qwen2.5-7B | 16.52 | 16.80 | 17.42 | 16.84 | 14.72 | 13.97 | 12.55 |
| Qwen3-8B | 12.26 | 12.09 | 11.38 | 11.07 | 10.93 | 11.61 | 10.70 |
| Model | Variant | variant | full | ||
| Mistral-7B | 0.2 | fixed | 7.786 | 7.33 | |
| Llama-3-8B | 0.2 | fixed | 10.975 | 10.44 | |
| Qwen3-8B | 0.2 | fixed | 11.100 | 10.70 | |
| Mistral-7B | 0.6 | with QKV metric | 8.456 | 5.44 | |
| Mistral-7B | 0.2 | with QKV metric | 303.397 | 7.33 | |
| Llama-3-8B | 0.6 | with QKV metric | 53.677 | 6.64 |
| Model | Variant | Full | |
| Llama-3-8B | 6.655 | 6.613 | 6.607 |
| Qwen3-8B | 7.855 | 7.764 | 7.759 |
| Qwen2.5-7B | 7.322 | 7.312 | 7.308 |
| Model | ppl | avg6 | GSM8K | ppl | avg6 | GSM8K | ppl | avg6 | GSM8K | ppl | avg6 | GSM8K |
| GPT-J-6B | 9.04 | 0.613 | 0.010 | 9.30 | 0.598 | 0.005 | 10.03 | 0.564 | 0.025 | 13.60 | 0.448 | n/a |
| Llama-2-13B | 4.98 | 0.715 | 0.190 | 5.11 | 0.703 | 0.115 | 5.52 | 0.660 | 0.030 | 7.52 | 0.497 | n/a |
| Mistral-7B | 5.36 | 0.733 | 0.235 | 5.44 | 0.726 | 0.240 | 5.70 | 0.689 | 0.105 | 7.33 | 0.512 | 0.015 |
| Llama-3-8B | 6.36 | 0.729 | 0.395 | 6.64 | 0.716 | 0.290 | 7.31 | 0.668 | 0.100 | 10.44 | 0.488 | 0.020 |
| Qwen2.5-7B | 6.99 | 0.715 | 0.780 | 7.20 | 0.689 | 0.715 | 7.84 | 0.659 | 0.445 | 12.55 | 0.405 | 0.015 |
| Model | (recipe) | |||
| Llama-3-8B | .6 | 6.64 | 6.63 ( ) | 6.54 ( ) |
| Llama-3-8B | .2 | 10.44 | 10.59 ( ) | 12.21 ( ) |
| Mistral-7B | .6 | 5.44 | 5.44 ( ) | 5.45 ( ) |
| Mistral-7B | .2 | 7.33 | 7.43 ( ) | 8.19 ( ) |
| Qwen3-8B | .6 | 7.36 | 7.39 ( ) | 7.58 ( ) |
| Qwen3-8B | .2 | 10.70 | 10.99 ( ) | 13.36 ( ) |
| Model | QKV target | layers escalated | reaching |
| Mistral-7B | no QKV map (final) | 4/32 | 0 |
| Mistral-7B | map | 5/32 | 1 |
| Llama-3-8B | no QKV map (final) | 2/32 | 0 |
| Llama-3-8B | map | 2/32 | 0 |
| Qwen2.5-7B | no QKV map (final) | 9/28 | 3 |
| Qwen2.5-7B | map | 5/28 | 0 |
| Model | 256 windows | 128 windows | ||
| Llama-3-8B | 0.6 | 6.640 | 6.660 | |
| Llama-3-8B | 0.2 | 10.440 | 10.821 | |
| Mistral-7B | 0.6 | 5.440 | 5.447 | |
| Mistral-7B | 0.2 | 7.330 | 7.499 | |
| Qwen2.5-7B | 0.6 | 7.200 | 7.192 | -0.1% |
| Qwen2.5-7B | 0.2 | 12.550 | 43.901 | +249.8% |
| Model | Unstructured core | 2:4 core | ||
| Mistral-7B | 0.6 | 5.44 | 5.47 | |
| Mistral-7B | 0.4 | 5.70 | 5.76 | |
| Mistral-7B | 0.2 | 7.33 | 8.64 | |
| Llama-3-8B | 0.6 | 6.64 | 6.72 | |
| Llama-3-8B | 0.4 | 7.31 | 7.46 | |
| Llama-3-8B | 0.2 | 10.44 | 12.16 |
| Model | Method | Calibration | wt2 | PTB | C4 | |
| Mistral-7B | 0.6 | SVD-LLM | WikiText-2 | 5.92 | 48.17 | 10.76 |
| Mistral-7B | 0.6 | SVD-LLM | C4 | 6.59 | 46.76 | 9.74 |
| Mistral-7B | 0.6 | MPIFA | WikiText-2 | 5.63 | 41.47 | 9.78 |
| Mistral-7B | 0.6 | MPIFA | C4 | 6.01 | 41.09 | 9.17 |
| Mistral-7B | 0.6 | Basis | WikiText-2 | 5.75 | 46.79 | 10.11 |
| Mistral-7B | 0.6 | Basis | C4 | 6.21 | 46.32 | 9.38 |
| Model | WT2 ( ) | WT2 ( ) | avg6 ( ) |
| Mistral-7B | |||
| Llama-3-8B | |||
| Qwen2.5-7B | |||
| Qwen3-8B |