FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
Authors: FangZhi Zhong, Xuerui Qiu, Yuqi Pan, Ya Liu, Shaowei Gu, Bo Xu, Guoqi Li
Organizations: Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Zhongguancun Academy · Shanghai Jiao Tong University
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at 2.9× input compression, including tool observations, versus 57.5 for Glyph at 3.0× input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a 2.79× online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
Figures & tables
Figure 1: Breaking the compression–performance trade-off through selective enhancement. (a) Fixed-resolution VTC ties local legibility to whole-page token cost. (b) FocusVTC retains a compressed global view and selectively enhances relevant regions. (c) On RULER v1, it sustains high scores at low initial DPI, while Glyph needs higher DPI and more tokens for comparable performance. (d) These gains coexist with preserved general capabilities across six benchmarks.
Figure 2: Overview of FocusVTC data construction and two-stage training. After calibrating font and point size, text contexts are rendered at multiple DPIs and annotated with REL-CoT. REL-SFT teaches low-DPI evidence localization, while GRPO trains the model to iteratively acquire and use high-resolution observations through selective region enhancement.
Figure 3: Font and point-size selection at 72 DPI. (a) Random-text CER across point sizes, with one-standard-error bars; references mark 5% CER and 9 pt. (b) 9 pt cost–CER comparison; the star indicates the preferred direction. (c) Confusable-unit errors, DejaVu Sans minus Verdana, with 95% paired-bootstrap intervals.
Model / stage
Input
Single-doc QA
Multi-doc QA
Summarization
Few-shot
Synthetic
Overall
QP
MF-En
MF-Zh
DuR
2Wiki
MNews
QMSum
SAMSum
Trivia
PR-En
PR-Zh
Avg
LLaMA-3.1-8B ( Grattafiori et al., 2024 )
Text
44.56
44.61
41.26
19.06
46.67
25.30
23.28
35.46
89.12
99.50
62.20
48.27
Qwen2.5-7B ( Yang et al., 2025b )
Text
45.29
43.44
42.12
16.55
40.51
24.94
22.95
34.59
86.93
100.00
98.50
50.53
Qwen3-8B ( Yang et al., 2025a )
Text
44.67
47.73
45.21
21.19
73.92
21.30
19.60
35.01
87.98
97.26
100.00
53.99
GLM-4-9B ( Team GLM et al., 2024 )
Text
43.75
45.21
41.23
20.79
50.89
24.82
22.84
35.84
90.07
99.50
100.00
52.27
Qwen3.5-9B ( Qwen Team, 2026 )
Text
47.23
49.32
56.58
25.25
61.82
23.79
22.28
38.02
90.16
100.00
100.00
55.86
Table 1: LongBench task scores (%), grouped by task family. Avg is the unweighted arithmetic mean of the 11 displayed non-code task scores. Vision denotes 72-DPI pages.
Figure 4: MRCR results: (a) two needles; (b) four needles; (c) eight needles; (d) prompt-plus-observation compression. Text is dashed/hollow; VTC is solid/filled. Context axes use bin upper bounds on a base-two scale. Compression divides Text tokens by the initial prompt plus tool observations. Full scores: Appendix C.3 .
Model / stage
Input
Retrieval
Reasoning
Memory
8
16
32
64
Avg
8
16
32
64
Avg
8
16
32
64
Avg
Qwen3-VL-8B ( Bai et al., 2025 )
Vision
90.50
77.90
77.01
76.36
80.44
18.83
12.16
11.11
1.52
10.91
21.08
20.18
24.03
20.00
21.32
Qwen3.5-9B ( Qwen Team, 2026 )
Vision
89.14
78.45
79.21
78.23
81.26
62.97
55.41
34.26
13.64
41.57
32.50
22.75
18.42
10.10
20.94
GLM-4.1V-9B ( V Team et al., 2025 )
Vision
86.20
85.25
82.76
58.01
78.06
31.59
10.14
22.22
9.09
18.26
25.10
21.89
13.16
7.07
16.81
Glyph ( Cheng et al., 2026 )
Vision
92.08
87.85
80.51
77.16
84.40
22.18
6.08
15.74
10.61
13.65
22.11
17.11
21.47
19.00
19.92
FocusVTC (Ours)
Vision
98.39
93.37
85.34
86.89
91.00
45.24
43.24
35.19
21.33
36.25
32.98
22.07
24.64
25.64
26.33
Table 2: VTCBench performance (%) across Retrieval, Reasoning, and Memory. Length labels are bin upper bounds (K tokens). Each Avg is the unweighted arithmetic mean of the four displayed length-bin scores. Vision denotes 72-DPI pages.
Model
REL-SFT
GRPO
Tools
RULER
LongBench
MRCR (needles)
VTCBench
v1
v2
Avg
2
4
8
Retrieval
Reasoning
Memory
FocusVTC (Ours)
✓
✓
✓
87.38
75.94
56.40
60.76
45.21
30.71
91.00
36.25
26.33
FocusVTC (Verdana) (Ours)
✓
✓
✓
85.13
75.49
54.93
60.26
43.11
27.34
89.21
35.93
24.23
FocusVTC w/o tools (Ours)
✓
✓
–
36.60
44.82
36.82
52.00
37.65
23.50
67.25
3.11
22.95
FocusVTC w/o SFT (Ours)
–
✓
✓
73.21
62.77
49.30
57.05
39.36
26.66
87.17
38.15
23.56
FocusVTC w/o GRPO (Ours)
✓
–
–
32.28
43.58
37.86
44.94
28.83
21.30
80.27
20.37
17.68
Table 3: Training stages and tool access at 72 DPI (scores in %). Best and second-best scores are bold and underlined, respectively. +tools uses tool access without additional training. Detailed results are provided in Appendices C.3 and C.4 .
Figure 5: Training behavior and checkpoint scores. (a) Tool calls; (b) response length; (c) grounding IoU; (d) benchmark scores at 0/50/100/150 steps. S1–S4 mark exploration, frequent tool use, fewer calls with longer responses, and shorter responses with higher IoU.
Figure 6: RULER results across DPI: (a) v1 scores; (b) v2 scores; (c) v1 input token costs; (d) v2 token costs. FocusVTC shows prompt and prompt-plus-observation costs. Glyph uses prompt tokens for v1 and v2. Dashed references show Text prompts. Token accounting: Appendix C.1 ; task-level scores at 72 DPI: Appendix C.4 .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Setting
Font family
DejaVu Sans (selected from 15 candidates)
Font size
9 pt (selected from 4, 5, 6, 7, 8, 9, 11, 13 pt)
Initial training DPI
48, 60, 72, 84, 96, 120, 144 for REL-SFT; 72, 96, 144 for GRPO
Default evaluation DPI
72, except for the resolution sweep
Enhancement source DPI
144, with aligned page layout and coordinates
Font-selection page layout and preprocessing
Appendix
Table 4: Rendering settings. The first block lists the selected font and resolution settings; page-layout and image-processor values in the second block apply to the font-selection diagnostics.
Font
Family
CER (%)
Threshold pt [95% CI]
Pages
Visual tokens
Liberation Sans Narrow
Sans
17.53
11.69 [11.41, 11.91]
12.18
6,016.9
Times New Roman
Serif
13.15
10.87 [10.76, 11.01]
13.54
6,688.8
Liberation Serif
Serif
12.94
10.67 [10.56, 10.80]
13.54
6,688.8
FreeSans
Sans
9.30
10.60 [10.44, 10.77]
14.50
7,163.0
Georgia
Serif
10.76
10.73 [10.54, 10.90]
14.82
7,321.1
Arial
Sans
8.39
10.39 [10.19, 10.59]
14.86
7,340.8
Appendix
Table 5: All font candidates at 72 DPI, sorted by visual-token cost. CER uses the 400-passage 9 pt set; threshold sizes and 95% intervals use the separate 30-passage sweep. Pagination and visual tokens are means over the same 50 long passages. Bold identifies the selected font.
Model / font
Single-1
Single-2
Single-3
MKey-1
MKey-2
MKey-3
MValue
MQuery
QA-1
QA-2
FocusVTC (Ours)
98.00
100.00
72.00
97.00
96.00
46.00
96.25
98.50
89.00
81.00
FocusVTC (Verdana) (Ours)
98.00
99.00
63.00
94.00
96.00
40.00
95.80
98.50
87.00
80.00
Appendix
Table 6: FocusVTC font comparison on RULER v1 at 72 DPI (scores in %).
Figure 7: REL-CoT source families (left) and the Multi-hop breakdown (right), before resolution expansion. Percentages use each panel’s total; the right panel combines long and original variants of each dataset.
Global batch size
Training steps
Learning rate
Weight decay
LR schedule
Warm-up steps
Visual encoder
Max. packed seq. length
8
7,000
10−6
0.01
Cosine
325
Frozen
32,768
Appendix
Table 7: Stage 1: REL-SFT training hyperparameters.
si=1 : the area factor is zero, so this call contributes zero to the matching sum.
Repeated requests for the same evidence
Each annotated evidence region can be matched at most once.
Excess or malformed tool calls
Both set Rfmt=0 . For N>E , Ecall=E/N<1 ; malformed calls count in N but are excluded from matching.
Initial view at 144 DPI
g(144)=0 , giving R=0.8Racc+0.2Rfmt .
Appendix
Table 9: Representative cases under the reward function. Unless otherwise stated, terminal answer formatting is valid.
Glyph ( Cheng et al., 2026 )
FocusVTC (Ours)
Benchmark
DPI
P
Text/P
P
O
P+O
Text/ (P+O)
RULER v1
48
1,491
5.6 ×
1,218
1,768
2,986
2.8 ×
60
2,061
4.1 ×
1,655
1,392
3,047
2.8 ×
72
2,765
3.0 ×
2,154
723
2,877
2.9 ×
84
3,517
2.4 ×
2,729
698
3,427
2.5 ×
96
4,522
1.9 ×
3,544
681
4,225
2.0 ×
Appendix
Table 10: Token costs and compression. P and O denote initial prompt and extra-observation tokens. Text references are 8,400/9,201/11,313 tokens for RULER v1/v2 and LongBench.
Model
Latency (s)
Initial prefill (s)
Tool-round prefill (s/sample)
Text (Qwen3.5-9B)
187.09
6.42
–
FocusVTC (Ours)
67.06
4.14
1.34
Appendix
Table 11: End-to-end latency on 300 paired MRCR four-needle examples with 64K–128K contexts. Values exclude recorded page-rendering time but include all online inference and tool rounds. Prefill timings are measured separately; tool-round prefill is cumulative per sample. Lower is better.
Model / stage
Input
Single-doc QA
Multi-doc QA
Summarization
Few-shot
Synthetic
Overall
QP
MF-En
MF-Zh
DuR
2Wiki
MNews
QMSum
SAMSum
Trivia
PR-En
PR-Zh
Avg
FocusVTC w/o SFT (Ours)
Vision
43.77
46.73
53.80
26.74
61.26
24.07
23.95
31.56
85.31
72.32
72.74
49.30
FocusVTC w/o GRPO (Ours)
Vision
34.52
39.94
31.77
14.73
56.54
17.47
16.69
31.64
89.23
67.12
16.85
37.86
FocusVTC w/o GRPO+tools (Ours)
Vision
16.46
32.70
23.16
8.42
49.70
2.96
3.79
3.88
56.42
65.19
16.83
25.41
FocusVTC-50 (Ours)
Vision
47.61
41.73
49.21
29.80
52.73
21.47
12.15
30.78
87.21
94.05
93.66
50.95
FocusVTC-100 (Ours)
Vision
46.90
45.06
54.22
30.10
61.14
22.47
20.06
32.76
90.10
95.73
96.00
54.05
Appendix
Table 12: Detailed LongBench scores (%) for training-stage variants. Avg is the unweighted arithmetic mean of the 11 displayed non-code task scores. The final FocusVTC row is included as reference. Vision denotes 72-DPI pages. +tools uses tool access without additional training.
Model / stage
Input
Retrieval
Reasoning
Memory
8
16
32
64
Avg
8
16
32
64
Avg
8
16
32
64
Avg
FocusVTC w/o SFT (Ours)
Vision
99.83
85.55
82.22
81.07
87.17
50.88
49.09
33.96
18.65
38.15
31.24
22.21
21.57
19.20
23.56
FocusVTC w/o GRPO (Ours)
Vision
87.78
79.56
80.21
73.52
80.27
40.17
25.68
12.04
3.57
20.37
12.12
17.98
20.60
20.00
17.68
FocusVTC w/o GRPO+tools (Ours)
Vision
66.28
59.67
51.72
39.34
54.25
7.82
1.35
0.93
0.00
2.53
7.50
4.23
6.18
9.79
6.93
FocusVTC-50 (Ours)
Vision
96.96
80.65
81.97
65.56
81.29
38.51
29.37
14.22
7.00
22.28
17.63
13.65
18.06
24.99
18.58
FocusVTC-100 (Ours)
Vision
97.94
87.85
85.34
77.05
87.05
39.53
29.73
16.67
11.21
24.29
25.53
17.84
20.85
25.64
22.47
Appendix
Table 13: Detailed VTCBench results (%) for training-stage variants. Length labels are bin upper bounds (K tokens). Each Avg is the unweighted arithmetic mean of the four displayed length-bin scores. The final FocusVTC row is included as reference. Vision denotes 72-DPI pages. +tools uses tool access without additional training.
Model / stage
Input
2 needles
4 needles
8 needles
8
16
32
64
128
256
Avg
8
16
32
64
128
256
Avg
8
16
32
64
128
256
Avg
LLaMA-3.1-8B ( Grattafiori et al., 2024 )
Text
54.27
53.21
51.05
29.81
24.98
20.90
39.04
33.42
25.97
22.73
26.97
12.68
6.00
21.30
23.80
17.69
19.85
17.72
11.79
7.80
16.44
Qwen2.5-7B ( Yang et al., 2025b )
Text
45.92
51.07
46.97
34.67
37.57
37.60
42.30
25.96
20.13
19.93
24.25
17.29
12.30
19.98
17.64
19.48
12.41
14.80
14.24
13.70
15.38
Qwen3-8B ( Yang et al., 2025a )
Text
58.95
41.18
36.18
24.99
20.89
17.50
33.28
29.34
22.67
20.34
23.63
19.11
15.50
21.77
18.75
19.69
16.81
17.86
15.00
12.60
16.79
GLM-4-9B ( Team GLM et al., 2024 )
Text
39.77
15.87
18.42
18.63
18.42
18.20
21.55
15.17
13.78
9.18
20.27
15.05
11.20
14.11
14.55
9.65
9.34
9.47
8.97
8.50
10.08
Qwen3.5-9B ( Qwen Team, 2026 )
Text
54.96
51.59
48.59
43.44
41.48
39.08
46.52
36.72
35.02
34.55
25.44
25.82
21.66
29.87
23.04
20.50
21.53
12.68
21.13
12.48
18.56
Appendix
Table 14: MRCR performance (%) for two, four, and eight needles. Length labels are bin upper bounds (K tokens). Avg is the unweighted arithmetic mean of the six length-bin scores. Vision denotes 72-DPI pages. +tools uses tool access without additional training.
Model
Single-1
Single-2
Single-3
MKey-1
MKey-2
MKey-3
MValue
MQuery
QA-1
QA-2
Avg
Qwen3.5-9B
54.00
48.00
0.00
43.00
40.00
1.00
37.25
42.75
62.00
46.00
37.40
GLM-4.1V-9B
61.00
26.00
0.00
30.00
19.00
0.00
19.50
33.25
50.00
48.00
28.68
Glyph
74.00
76.00
44.00
69.00
37.00
2.00
70.75
74.50
64.00
64.00
57.53
FocusVTC w/o GRPO (Ours)
50.00
39.00
0.00
29.00
37.00
0.00
24.25
24.50
65.00
54.00
32.28
FocusVTC w/o GRPO+tools (Ours)
43.00
29.00
1.00
24.00
37.00
1.00
23.75
23.00
22.00
26.00
22.98
FocusVTC-50 (Ours)
82.99
82.73
46.03
73.50
66.15
37.59
54.36
71.10
77.27
68.58
66.03
Appendix
Table 15: RULER v1 task scores (%) at 72 DPI. Avg is the unweighted arithmetic mean of the 10 displayed task scores. Best and second-best scores are bold and underlined, respectively. Qwen3.5-9B+tools uses tool access without additional training.
Model
MK-NIAH
MV-NIAH
QA
Overall
Basic
Easy
Medium
Hard
Basic
Easy
Medium
Hard
Basic
Easy
Medium
Hard
Avg
Qwen3.5-9B
19.00
79.44
70.00
46.00
10.75
8.65
23.36
23.45
69.00
34.00
71.00
81.00
44.64
GLM-4.1V-9B
8.00
51.12
29.00
25.00
3.50
9.67
10.64
18.24
53.00
32.00
53.58
53.50
28.94
Glyph
16.00
67.15
45.00
43.00
9.25
5.72
17.02
34.05
86.00
93.00
78.35
80.36
47.91
FocusVTC w/o GRPO (Ours)
18.00
59.96
51.00
39.00
7.25
7.05
16.79
27.24
74.00
83.00
68.50
71.17
43.58
FocusVTC w/o GRPO+tools (Ours)
20.00
71.32
45.00
33.00
10.75
7.03
13.04
29.78
72.00
36.00
63.27
71.68
39.41
Appendix
Table 16: RULER v2 task scores (%) at 72 DPI. Avg is the unweighted arithmetic mean of the 12 displayed task scores. Best and second-best scores are bold and underlined, respectively. Qwen3.5-9B+tools uses tool access without additional training.
Model
OCRBench
DocVQA
MMMU
MME
ChartQA
InfoVQA
Qwen3.5-9B ( Qwen Team, 2026 )
851
92.38
65.12
2424.02
85.96
74.76
Glyph ( Cheng et al., 2026 )
799
91.75
57.67
2253.13
72.76
71.89
GLM-4.1V-9B ( V Team et al., 2025 )
820
91.34
64.33
2392.53
70.76
70.52
FocusVTC (Ours)
860
92.43
66.73
2457.62
86.28
76.20
Model
VRAG
VH
NIAH image
NIAH text
Summ
DocQA
Qwen3.5-9B ( Qwen Team, 2026 )
61.12
56.10
45.04
71.82
29.19
68.96
Appendix
Table 17: General multimodal scores from Figure 1 (d) (top) and MMLongBench results (bottom). DocVQA and InfoVQA use ANLS; other entries retain their benchmark scales.
Figure 9: A factual evidence chain in REL-CoT. The question and target are adapted; source pages and boxes are unchanged. (A) A factual question over 22 pages (72-DPI thumbnails). (B) 144-DPI excerpts linking Merlin, Arthur, and Sir Ector; box coordinates are normalized to [0,1000] . (C) Initial question analysis followed by sentence-embedded page–box citations and a short, directly verifiable answer.
Figure 10: An adapted MuSiQue illustration of selective enhancement. The question, source pages, reasoning excerpts, and answer are retained from the recorded Verdana example; tool-call boxes and observations are adapted for illustration. Enhance_Region reads local passages on pages 3 and 6 from aligned 144-DPI sources, using original-page coordinates in [0,1000] . Insets enlarge text within these enhanced regions.
Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolution, making them indistinguishable. To address this, we propose LensVLM, an inference framework and post-training recipe that enables VLMs to scan compressed images, then selectively expand only the relevant images to their uncompressed form via learned tools. Building on Qwen3.5-9B-Base, LensVLM maintains accuracy comparable to the full-text upper bound at 4.3× effective compression and outperforms retrieval-based, text- and visual-compression baselines up to 10.1× effective compression across seven text QA benchmarks. LensVLM also generalizes to multimodal document and code understanding tasks, with the accuracy gain over baselines growing as compression increases. Our analysis validates this approach: training makes visual compression robust to rendering choices, and as compression grows the model increasingly relies on expanded content rather than unreliable visual reading. The analysis also yields practical tool-choice guidance: text expansion is preferable for rendered text, while high-resolution image expansion suits native documents whose layout cues carry task-relevant information.
Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions either resort to aggressive token pruning, risking irreversible information loss, or adopt efficient but less precise architectures, while largely ignoring the equally vital textual component. We introduce VLZip, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer. At its core, VLZip hierarchically distills visual and textual segments into compact, layer-specific "soft prefixes" and injects them into each decoder layer's hidden states, drastically shortening the attention sequence while preserving fine-grained global context. To address deficient evaluations in the field, we also introduce LongVLBench, a new benchmark derived from video narratives that demands holistic, narrative-level reasoning. Extensive experiments show VLZip achieves leading performance on long-context multimodal reasoning, enabling training up to 120K tokens, a 6x increase over the baseline, and inference beyond 280K tokens with significantly reduced memory, while demonstrating the memory scalability to handle up to 2M tokens. By excelling at extreme context lengths where existing methods collapse, VLZip establishes an efficient and powerful new standard for long-context multimodal AI. Code is available at https://github.com/ShareLab-SII/VLZip.
Yuqi Zhang, Cheng Chen, Yuyu Guo +6
Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University, China · Shanghai Innovation Institute, China · Ant Group, China
Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing 3--20× fewer decoder tokens than subword tokenization. Yet token savings do not translate predictably into downstream utility: on some tasks the visual path matches or exceeds the text path, on others it collapses, and the compression ratio itself does not predict which regime will occur. The missing quantity is therefore not another summary of efficiency, but a principled measure of task-relevant information loss induced by visual encoding. We address this problem by formulating VTC in the language of measure transport. Treating text and visual tokens as empirical probability measures, we show that the ViT patch encoder induces a push-forward map whose transport cost decomposes into a precision cost from within-patch aggregation and a coverage cost from cross-patch fragmentation. Both terms are estimable from downstream-label-free probes. This formulation yields two operational consequences: a downstream-label-free routing criterion that selects whether to use the visual path for a given input or benchmark instance, and a transport-informed foveation mechanism that re-encodes high-cost regions at higher resolution. Across 24 NLP datasets at Qwen3-4B, our label-free rule matches the per-dataset oracle on 17/24 datasets (70.8%), and improves the average task score by +3.3% with −10.3% average tokens relative to a pure-LLM.
Lv Tang, Tianyi Zheng, Yang Liu +2
University of Alberta1 · vivo Mobile Communication Co., Ltd2 · Tsinghua University3