Embedding tables are among the largest components of modern language models. Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it. We instead ask whether the coding geometry can itself be discovered. We introduce \emph{OrBIT}, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords. The global reconstruction residual then decides where the fixed coding budget is spent, while redundant overlapping charts let local errors compensate one another after gluing. Our theory shows how tight-chart geometry controls distortion, how the global residual directs sequential allocation, and how data-geometry-guided refinement improves the codec. The resulting orbit machinery is compiled away, leaving a compact decoder in which the learned structure governs what is stored, where capacity is allocated, and how local information is assembled globally. Across four LLM embedding tables, OrBIT achieves 37.9× compression on GPT-2 and over 23× on each 7B table relative to 16-bit storage, while delivering competitive rate-distortion performance against established quantization and low-rank baselines.
Figures & tables
Embedding table
OrBIT, T=7
PQ-q8
OPQ-q8
AQ-q8
GPT-2
324.15/0.7449
307.40/0.7467
313.96/0.7465
331.88/0.7582
Llama-2-7B
2739.97/0.9163
2703.56/0.9211
2762.55/0.9192
2911.54/0.9278
Mistral-7B
2739.97/0.9151
2703.56/0.9189
2762.55/0.9170
2911.54/0.9228
Mistral-7B-Instruct
2739.97/0.9205
2703.56/0.9176
2762.55/0.9157
2911.54/0.9205
Table 1: Representative end-to-end rate–distortion results. Each entry is bits/token / relative error in the original embedding space. The PQ-q8 and OPQ-q8 entries use m=16,b=8 ; AQ-q8 uses m=8,b=8 . All shared objects, including the PCA representation map, are charged to representative operating points with matched precision. Lower error and lower rate are better.
Model
Before
After
Reduction
GPT-2
0.4750
0.4555
4.11%
Llama-2-7B
0.3050
0.2966
2.76%
Mistral-7B
0.04777
0.04643
2.82%
Mistral-Inst.
0.04968
0.04830
2.77%
Table 2: Gluing-aware refinement. Rate is unchanged; lower absolute weighted PCA-space error is better.
Post-gluing error
Cancelled local error
Model
A=1
A=2
A=3
A=2
A=3
GPT-2
0.18893
0.18903
0.18897
98.85%
99.43%
Llama-2-7B
0.22596
0.22378
0.22057
81.27%
89.85%
Mistral-7B
0.03577
0.03492
0.03430
82.41%
90.66%
Mistral-Inst.
0.03968
0.03813
0.03814
77.48%
87.04%
Table 3: Matched redundancy ablation. The total chart-stage budget LT=48 is fixed. “Cancelled” denotes εinc2/εloc2 . GPT-2 uses 441.56 bits/token in every arm; the 7 B models use 2890.14 bits/token.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Original dim.
d
L
Mℓ
Anchor frac.
Max support
T
Bits/token
GPT-2
768
64
4
32
0.75%
32
7
324.1472
Llama-2-7B
4096
192
4
96
1.2%
96
7
2739.968
Mistral-7B
4096
192
4
96
1.2%
96
7
2739.968
Mistral-7B-Instruct
4096
192
4
96
1.2%
96
7
2739.968
Appendix
Table 4: Standard OrBIT configuration. The anchor fraction is measured relative to the D=10,000 -token table. The reported rate is the complete compiled rate, including the shared representation map.
Figure 1: Rate–distortion curves across the four embedding tables. Rate is measured in bits per token and distortion by relative reconstruction error in the original embedding space. Lower values are better on both axes.
Model
Chart
Orbit scaffold
Gaussian control
Best Gaussian
GPT-2
1
0.47290
0.49930±0.00641
0.48787
2
0.47890
0.50236±0.00688
0.48572
3
0.48054
0.50028±0.00525
0.49094
4
0.47782
0.50233±0.00623
0.49177
Llama-2-7B
1
0.47805
0.49885±0.00202
0.49489
2
0.47321
0.49905±0.00295
0.49120
Appendix
Table 5: Experiment 2: data-aligned orbit scaffold versus a rank-matched Gaussian control. Each Gaussian entry is the mean ± standard deviation over 20 independent random subspaces; “Best Gaussian” is the smallest residual among those 20 draws. All methods have rank r=0.75Mℓ , and evaluation uses the complete local table. Lower relative projection error is better.
Model
PCA before
PCA after
Reduction
End-to-end before/after
Accepted updates/280
GPT-2
0.47502
0.45549
4.11%
0.74570/0.74493
265
Llama-2-7B
0.30500
0.29659
2.76%
0.91865/0.91631
280
Mistral-7B
0.04777
0.04643
2.82%
0.91737/0.91518
280
Mistral-7B-Instruct
0.04968
0.04830
2.77%
0.92299/0.92048
280
Appendix
Table 6: Experiment 3: post-gluing refinement. The bitrate is identical before and after refinement. The PCA-space columns report the absolute weighted reconstruction error optimized by refinement; the end-to-end column is measured after lifting to the original embedding space.
Model
A
εloc
εinc
Cancelled
Global residual
End-to-end rel. error
GPT-2
1
0.18893
3.38×10−7
0.00%
0.18893
0.73763
2
2.48802
2.47362
98.85%
0.18903
0.73764
3
4.32155
4.30914
99.43%
0.18897
0.73763
Llama-2-7B
1
0.22596
1.51×10−7
0.00%
0.22596
0.89909
2
0.73135
0.65933
81.27%
0.22378
0.89863
3
1.19937
1.13690
89.85%
0.22057
0.89796
Appendix
Table 7: Experiment 4: matched-capacity redundancy audit. “Cancelled” is εinc2/εloc2 . The global residual is the weighted reduced-space norm entering the tight-frame identity. All three values of A have identical bitrate within each model.
Model
Explicit bits/token
Compiled bits/token
Saved bits/token
Storage reduction
Reconstruction gap
GPT-2
355.5468
324.1472
31.3996
8.83%
0
Llama-2-7B
2889.1328
2739.9680
149.1648
5.16%
0
Mistral-7B
2889.1328
2739.9680
149.1648
5.16%
0
Mistral-7B-Instruct
2889.1328
2739.9680
149.1648
5.16%
0
Appendix
Table 8: Experiment 5: compilation audit. “Explicit” retains the orbit/scaffold construction state in addition to the compiled payload. “Compiled” retains only the objects required by the deployed decoder. Reconstruction is identical in the two cases.
Model
Anchor frac.
Anchors
Rank / Mℓ
Mean εproj
All charts accepted
GPT-2
0.50%
50
32/32
6.31×10−7
Yes
0.64%
64
32/32
6.19×10−7
Yes
0.75%
75
32/32
6.74×10−7
Yes
1.00%
100
32/32
5.88×10−7
Yes
Llama-2-7B
0.50%
50
50/96
2.3794×10−1
No
0.64%
64
64/96
1.9669×10−1
No
Appendix
Table 9: Experiment 6: anchor-budget diagnostic. The rank column is the scaffold rank in a chart of dimension Mℓ . The projection error is averaged over the four charts and is always evaluated on the complete local table.
Model
PCA-space rel. error
End-to-end rel. error
Mean cosine
Recall@10
Bits/token
GPT-2
0.26113±0.00063
0.74506±0.00031
0.66730±0.00084
0.29727±0.01412
324.1472±0
Llama-2-7B
0.60226±0.01186
0.91820±0.00158
0.39200±0.00323
0.33103±0.00466
2739.968±0
Mistral-7B
0.57864±0.00966
0.91598±0.00120
0.39796±0.00303
0.36715±0.00710
2739.968±0
Mistral-7B-Instruct
0.62136±0.00638
0.92099±0.00105
0.37600±0.00228
0.32091±0.00454
2739.968±0
Appendix
Table 10: Experiment 7: robustness of the final T=7 headline operating point over seeds {13,29,47} . Values are mean ± standard deviation.
High-dimensional language-model embeddings increase storage and search costs, while supervised compressors can overfit when relevance labels are scarce. We present DIVE (Dimensionality reduction with Implicit View Ensembles), a residual compression adapter codesigned with a self-limiting hinge loss, geometry distillation, and head-wise NT-Xent over implicit coordinate views. The hinge stops updating satisfied ranking constraints, while the dense objectives stabilize the compressed representation; only the first head is retained at inference. Under query-disjoint evaluation with two LLM2Vec backbones, five BEIR benchmarks, 128d and 256d outputs, and six baselines, DIVE is the strongest adapter on all five primary benchmarks. It also outperforms PCA and an autoencoder in comparisons against unsupervised compressors.
Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically with input length. Token compression aims to address this challenge by reducing the number of tokens representing inputs. However, existing prompt-compression approaches primarily operate in token space and overlook inefficiencies in the latent embedding space. In this paper, we propose K-Token Merging, a latent-space compression framework that merges each contiguous block of K token embeddings into a single embedding via a lightweight encoder. The compressed sequence is processed by a LoRA-adapted LLM, while generation remains in the original vocabulary. Experiments on structural reasoning (Textualized Tree), sentiment classification (Amazon Reviews), and code editing (CommitPackFT) show that K-Token Merging lies on the Pareto frontier of performance vs. compression, achieving up to 75% input length reduction with minimal performance degradation. Code is available at https://github.com/shsjxzh/K-Token-Merging.
Zihao Xu, John Harvill, Ziwei Fan +3
♢Rutgers University · ♡Mistral AI · ♠AWS AI Labs +3
Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.
Jiangfeng Chen, Xinyu Wang, Tianshuo Yan +4
University of Manitoba · McGill University · Simpleway +2