Organizations: State Key Laboratory for Novel Software Technology, Nanjing University · National University of Singapore · Nanjing University of Posts and Telecommunications
Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long documents, such mean pooling can dilute salient semantic information with abundant redundant or weakly informative content. To this end, we propose SCSP, a training-free framework that leverages semantic compression for informative token selection in long-context embedding. Specifically, SCSP first partitions a document into sentence-aware chunks and appends a semantic compression prompt to each chunk. A prompt-isolated attention mask preserves information flow among document tokens while restricting each prompt to its corresponding local context. We then use the attention patterns elicited by these prompts to estimate token importance, select informative tokens, and aggregate their intermediate-layer representations into the final embedding. Extensive experiments on long-context embedding benchmarks demonstrate that SCSP can be integrated into both zero-shot and fine-tuned models in a plug-and-play manner, consistently improving their performance.
Figures & tables
Figure 1: Illustration of the SCSP method.
CXT Len
Method
QMSum
2WikiMQA
SumFD
NQA
MultiFieldQA
Avg
Times
512
PromptEOL †
4.57
7.16
6.07
2.80
21.72
8.46
–
TP w. PromptEOL †
5.44
6.51
5.52
1.72
17.58
7.35
–
TP w. Mean †
11.97
18.62
38.50
3.08
44.80
23.39
–
Vanilla Mean ‡
15.01
16.07
40.39
2.56
48.61
24.53
1.00×
Vanilla Mean + SCSP ( Ours )
14.07
21.21
45.63
3.27
52.06
27.25 (+2.72)
1.19×
Echo Mean ‡
14.34
25.63
32.36
7.06
60.37
27.95
1.65×
Table 1: NDCG@10 (in percentage) on five datasets using Mistral-7B-Instruct-v0.3. We report the context length of 512 and an extended length of 8192. For the first four datasets that were also evaluated in HTP, † denotes the results reported in HTP, while ‡ denotes our reproduced results, which closely match the reported results.
CXT Len
Method
QMSum
2WikiMQA
SumFD
NQA
MultiFieldQA
Avg
Times
512
PromptEOL †
10.17
15.20
11.39
5.22
46.70
17.73
–
TP w. PromptEOL †
8.33
8.29
10.97
1.91
25.65
11.03
–
TP w. Mean †
7.13
5.29
6.95
1.57
17.72
7.73
–
Vanilla Mean ‡
12.21
18.09
35.92
2.82
45.39
22.89
1.00×
Vanilla Mean + SCSP ( Ours )
11.60
26.08
48.57
3.48
49.86
27.92 (+5.03)
1.34×
Echo Mean ‡
13.86
22.85
36.35
7.56
67.83
29.69
1.90×
Table 2: NDCG@10 (in percentage) on five datasets using LLaMA-3.1-8B-Instruct.
Method
QMSum
2WikiMQA
SumFD
NQA
MultiFieldQA
Avg
Vanilla Mean
26.67
25.92
58.04
5.39
46.18
32.44
Vanilla Mean + SCSP ( Ours )
28.89
34.46
70.58
10.62
51.19
39.15
Input Construction:
w/o Sentence-aware Chunking
28.65
33.47
69.17
8.27
51.13
38.14
w/o Chunk-wise Prompt
24.59
29.56
67.25
8.43
52.30
36.42
w/o Semantic Compression Prompt
26.17
29.08
59.50
8.59
48.55
34.38
Table 3: Ablation study results on Mistral-7B-Instruct-v0.3 with an 8,192-token context length.
Figure 2: The average effects of context length, chunk size, ratio α , and output layer on performance across five datasets are evaluated using Mistral-7B-Instruct-v0.3.
Backbone
Method
QMSum
2WikiMQA
SumFD
NQA
MultiFieldQA
Avg
Gemma2-9B
Vanilla Mean
30.78
33.82
70.55
7.95
49.99
38.62
Vanilla Mean + SCSP ( Ours )
32.72
43.56
73.61
11.55
53.35
42.96 (+4.34)
Qwen2.5-7B-Instruct
Vanilla Mean
25.55
22.43
47.49
6.33
50.16
30.39
Vanilla Mean + SCSP ( Ours )
26.66
25.90
54.77
7.90
45.52
32.15 (+1.76)
Table 4: Generalization across different LLM backbones on five datasets with an 8,192-token context length.
Method
Layer
Head
QMSum
2WikiMQA
SumFD
NQA
MultiFieldQA
Avg
Vanilla Mean
–
–
26.67
25.92
58.04
5.39
46.18
32.44
SCSP
Max
Max
28.89
34.46
70.58
10.62
51.19
39.15
Max
Mean
29.35
34.53
69.06
9.91
50.87
38.75
Mean
Max
27.58
33.20
67.38
10.69
50.57
37.89
Mean
Mean
25.91
30.30
63.08
10.10
49.15
35.71
Table 5: Comparison of different strategies for computing token importance scores on Mistral-7B-Instruct-v0.3 with an 8,192-token context length.
Models
Method
QMSum
2WikiMQA
SumFD
NQA
MultiFieldQA
Avg
GritLM
Vanilla Mean
19.98
27.63
29.48
5.89
72.46
31.09
Vanilla Mean + SCSP ( Ours )
24.34
37.27
57.41
14.74
78.31
42.42 (+11.33)
NV-EMBED
Vanilla Mean
24.87
34.77
67.11
24.77
74.61
45.23
Vanilla Mean + SCSP ( Ours )
28.12
34.99
72.55
29.15
76.63
48.29 (+3.06)
Table 6: Generalization of SCSP to fine-tuned embedding models on five datasets with an 8,192-token context length.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Prompt
QMSum
2WikiMQA
SumFD
NQA
MultiFieldQA
Avg
Brief summary: “
28.89
34.46
70.58
10.62
51.19
39.15
Brief summary:
28.62
34.14
67.59
10.03
51.54
38.38
This text means in one word: “
28.22
31.88
65.24
9.41
50.98
37.15
This text can be summarized as: “
28.01
33.72
65.92
9.5
52.23
37.87
Summary: “
28.42
35.05
70.07
10.05
52.15
39.15
Appendix
Table 7: Effects of semantic compression prompts on Mistral-7B-Instruct-v0.3 with an 8,192-token context length.
Dataset
Vanilla Mean
Vanilla Mean + SCSP ( Ours )
SumFD
57.68
70.24
Gov. Report
97.30
98.12
QMSUM
50.94
55.12
QASPER Title
37.27
46.40
QASPER Abstract
89.19
93.45
2WikiMQA
46.90
51.76
Appendix
Table 8: Results on the LOCOV1 dataset using Mistral-7B-Instruct-v0.3 with an 8,192-token context length.
Long-context large language models remain computationally expensive to run and often fail to reliably process very long inputs, which makes context compression an important component of many systems. Existing compression approaches typically rely on trained compressors, dense retrieval-style selection, or heuristic trimming, and they often struggle to jointly preserve task relevance, topic coverage, and cross-sentence coherence under a strict token budget. To address this, we propose a training-free and model-agnostic compression framework that selects a compact set of sentences guided by structural graph priors. Our method constructs a sparse hybrid sentence graph that combines mutual k-NN semantic edges with short-range sequential edges, extracts a topic skeleton via clustering, and ranks sentences using an interpretable score that integrates task relevance, cluster representativeness, bridge centrality, and a cycle coverage cue. A budgeted greedy selection with redundancy suppression then produces a readable compressed context in original order. Experimental results on four datasets show that our approach is competitive with strong extractive and abstractive baselines, demonstrating larger gains on long-document benchmarks.
Yitian Zhou, Chaoning Zhang, Jiaquan Zhang +6
School of Computer Science and Engineering University of Electronic Science and Technology of China · School of Information and Software Engineering, University of Electronic Science and Technology of China · Department of Computer Science and Engineering, Kyung Hee University +3
Large Language Models (LLMs) have demonstrated exceptional performance across diverse tasks. However, their deployment in long-context scenarios faces high computational overhead and information redundancy. While soft prompt compression has emerged as a promising way to mitigate these costs by compressing sequences into compact embeddings, existing paradigms remain fundamentally constrained by position bias: they primarily rely on learnable tokens insertion at fixed positions or group tokens according to their physical token layout, thereby inducing performance instability and semantic fragmentation. To overcome this bottleneck, we propose Semantic Consistency Context Compression (SeCo), a method that shifts context compression from position-driven to semantic-driven. Rather than constraint by physical token layout, SeCo dynamically anchors compression directly in the semantic space by selecting query-relevant tokens as semantic centers and aggregating remaining tokens via consistency-weighted merging. This design inherently preserves semantic consistency while eliminating position bias. Extensive experiments on 14 benchmarks across two backbone models demonstrate that SeCo consistently shows superiority in downstream tasks, inference latency, and out-of-domain robustness. The code is available at https://anonymous.4open.science/r/seco-EE5E.
Jiwei Tang, Zhijing Huang, Xinyu Zhang +5
Sun Yat-sen University · Hong Kong Polytechnic University · Beijing Normal–Hong Kong Baptist University
Long-context compression is essential for reducing the cost and latency of large language model inference. However, existing methods can fragment important evidence, require additional training or alignment, and often depend on the target model for effective compression. We introduce TopoCompress, a training-free and model-agnostic framework that compresses long contexts by selecting coherent semantic spans. TopoCompress first scores each span using dense and lexical query relevance together with semantic acceleration. It then constructs a hybrid graph that connects spans based on semantic similarity and sequential adjacency, and propagates the query-guided relevance scores over the graph. Across five long-context tasks-HotpotQA, 2WikiMQA, MuSiQue, Qasper, and MultiFieldQA-en-TopoCompress consistently outperforms strong compression baselines. Notably, TopoCompress achieves performance comparable to the strongest baseline while using a 4x smaller compression budget, and provides a 1.41x smaller compression time over the fastest baseline.