FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
Organizations: Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Zhongguancun Academy · Shanghai Jiao Tong University
Abstract
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at input compression, including tool observations, versus 57.5 for Glyph at input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
Figures & tables
| Model / stage | Input | Single-doc QA | Multi-doc QA | Summarization | Few-shot | Synthetic | Overall | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| QP | MF-En | MF-Zh | DuR | 2Wiki | MNews | QMSum | SAMSum | Trivia | PR-En | PR-Zh | Avg | ||
| LLaMA-3.1-8B ( Grattafiori et al., 2024 ) | Text | 44.56 | 44.61 | 41.26 | 19.06 | 46.67 | 25.30 | 23.28 | 35.46 | 89.12 | 99.50 | 62.20 | 48.27 |
| Qwen2.5-7B ( Yang et al., 2025b ) | Text | 45.29 | 43.44 | 42.12 | 16.55 | 40.51 | 24.94 | 22.95 | 34.59 | 86.93 | 100.00 | 98.50 | 50.53 |
| Qwen3-8B ( Yang et al., 2025a ) | Text | 44.67 | 47.73 | 45.21 | 21.19 | 73.92 | 21.30 | 19.60 | 35.01 | 87.98 | 97.26 | 100.00 | 53.99 |
| GLM-4-9B ( Team GLM et al., 2024 ) | Text | 43.75 | 45.21 | 41.23 | 20.79 | 50.89 | 24.82 | 22.84 | 35.84 | 90.07 | 99.50 | 100.00 | 52.27 |
| Qwen3.5-9B ( Qwen Team, 2026 ) | Text | 47.23 | 49.32 | 56.58 | 25.25 | 61.82 | 23.79 | 22.28 | 38.02 | 90.16 | 100.00 | 100.00 | 55.86 |
| Model / stage | Input | Retrieval | Reasoning | Memory | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 8 | 16 | 32 | 64 | Avg | 8 | 16 | 32 | 64 | Avg | 8 | 16 | 32 | 64 | Avg | ||
| Qwen3-VL-8B ( Bai et al., 2025 ) | Vision | 90.50 | 77.90 | 77.01 | 76.36 | 80.44 | 18.83 | 12.16 | 11.11 | 1.52 | 10.91 | 21.08 | 20.18 | 24.03 | 20.00 | 21.32 |
| Qwen3.5-9B ( Qwen Team, 2026 ) | Vision | 89.14 | 78.45 | 79.21 | 78.23 | 81.26 | 62.97 | 55.41 | 34.26 | 13.64 | 41.57 | 32.50 | 22.75 | 18.42 | 10.10 | 20.94 |
| GLM-4.1V-9B ( V Team et al., 2025 ) | Vision | 86.20 | 85.25 | 82.76 | 58.01 | 78.06 | 31.59 | 10.14 | 22.22 | 9.09 | 18.26 | 25.10 | 21.89 | 13.16 | 7.07 | 16.81 |
| Glyph ( Cheng et al., 2026 ) | Vision | 92.08 | 87.85 | 80.51 | 77.16 | 84.40 | 22.18 | 6.08 | 15.74 | 10.61 | 13.65 | 22.11 | 17.11 | 21.47 | 19.00 | 19.92 |
| FocusVTC (Ours) | Vision | 98.39 | 93.37 | 85.34 | 86.89 | 91.00 | 45.24 | 43.24 | 35.19 | 21.33 | 36.25 | 32.98 | 22.07 | 24.64 | 25.64 | 26.33 |
| Model | REL-SFT | GRPO | Tools | RULER | LongBench | MRCR (needles) | VTCBench | |||||
| v1 | v2 | Avg | 2 | 4 | 8 | Retrieval | Reasoning | Memory | ||||
| FocusVTC (Ours) | 87.38 | 75.94 | 56.40 | 60.76 | 45.21 | 30.71 | 91.00 | 36.25 | 26.33 | |||
| FocusVTC (Verdana) (Ours) | 85.13 | 75.49 | 54.93 | 60.26 | 43.11 | 27.34 | 89.21 | 35.93 | 24.23 | |||
| FocusVTC w/o tools (Ours) | – | 36.60 | 44.82 | 36.82 | 52.00 | 37.65 | 23.50 | 67.25 | 3.11 | 22.95 | ||
| FocusVTC w/o SFT (Ours) | – | 73.21 | 62.77 | 49.30 | 57.05 | 39.36 | 26.66 | 87.17 | 38.15 | 23.56 | ||
| FocusVTC w/o GRPO (Ours) | – | – | 32.28 | 43.58 | 37.86 | 44.94 | 28.83 | 21.30 | 80.27 | 20.37 | 17.68 | |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Setting |
|---|---|
| Font family | DejaVu Sans (selected from 15 candidates) |
| Font size | 9 pt (selected from 4, 5, 6, 7, 8, 9, 11, 13 pt) |
| Initial training DPI | 48, 60, 72, 84, 96, 120, 144 for REL-SFT; 72, 96, 144 for GRPO |
| Default evaluation DPI | 72, except for the resolution sweep |
| Enhancement source DPI | 144, with aligned page layout and coordinates |
| Font-selection page layout and preprocessing | |
| Font | Family | CER (%) | Threshold pt [95% CI] | Pages | Visual tokens |
|---|---|---|---|---|---|
| Liberation Sans Narrow | Sans | 17.53 | 11.69 [11.41, 11.91] | 12.18 | 6,016.9 |
| Times New Roman | Serif | 13.15 | 10.87 [10.76, 11.01] | 13.54 | 6,688.8 |
| Liberation Serif | Serif | 12.94 | 10.67 [10.56, 10.80] | 13.54 | 6,688.8 |
| FreeSans | Sans | 9.30 | 10.60 [10.44, 10.77] | 14.50 | 7,163.0 |
| Georgia | Serif | 10.76 | 10.73 [10.54, 10.90] | 14.82 | 7,321.1 |
| Arial | Sans | 8.39 | 10.39 [10.19, 10.59] | 14.86 | 7,340.8 |
| Model / font | Single-1 | Single-2 | Single-3 | MKey-1 | MKey-2 | MKey-3 | MValue | MQuery | QA-1 | QA-2 |
|---|---|---|---|---|---|---|---|---|---|---|
| FocusVTC (Ours) | 98.00 | 100.00 | 72.00 | 97.00 | 96.00 | 46.00 | 96.25 | 98.50 | 89.00 | 81.00 |
| FocusVTC (Verdana) (Ours) | 98.00 | 99.00 | 63.00 | 94.00 | 96.00 | 40.00 | 95.80 | 98.50 | 87.00 | 80.00 |
| Global batch size | Training steps | Learning rate | Weight decay | LR schedule | Warm-up steps | Visual encoder | Max. packed seq. length |
|---|---|---|---|---|---|---|---|
| 8 | 7,000 | 0.01 | Cosine | 325 | Frozen | 32,768 |
| Global batch | Train steps | LR | Weight decay | LR schedule | Warmup steps | Visual encoder | Max. prompt / response | Max. gen. / turn | Rollouts / prompt | PPO mini-batch | PPO micro- batch / GPU | Temp. / top- | Max. turns / tool calls | exponent | Clip |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 64 | 150 | 0.01 | Constant | 0 | Trainable | 8K / 10K | 2K | 8 | 64 | 1 | 1.0 / 1.0 | 9 / 8 | 2 | 0.2 |
| Situation | Consequence |
|---|---|
| Fully correct answer, no tool calls | and , giving . |
| Answer not fully correct | The tool-bonus term is gated off. |
| Full-page request | : the area factor is zero, so this call contributes zero to the matching sum. |
| Repeated requests for the same evidence | Each annotated evidence region can be matched at most once. |
| Excess or malformed tool calls | Both set . For , ; malformed calls count in but are excluded from matching. |
| Initial view at 144 DPI | , giving . |
| Glyph ( Cheng et al., 2026 ) | FocusVTC (Ours) | ||||||
| Benchmark | DPI | P | Text/P | Text/ | |||
| RULER v1 | 48 | 1,491 | 5.6 | 1,218 | 1,768 | 2,986 | 2.8 |
| 60 | 2,061 | 4.1 | 1,655 | 1,392 | 3,047 | 2.8 | |
| 72 | 2,765 | 3.0 | 2,154 | 723 | 2,877 | 2.9 | |
| 84 | 3,517 | 2.4 | 2,729 | 698 | 3,427 | 2.5 | |
| 96 | 4,522 | 1.9 | 3,544 | 681 | 4,225 | 2.0 | |
| Model | Latency (s) | Initial prefill (s) | Tool-round prefill (s/sample) |
|---|---|---|---|
| Text (Qwen3.5-9B) | 187.09 | 6.42 | – |
| FocusVTC (Ours) | 67.06 | 4.14 | 1.34 |
| Model / stage | Input | Single-doc QA | Multi-doc QA | Summarization | Few-shot | Synthetic | Overall | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| QP | MF-En | MF-Zh | DuR | 2Wiki | MNews | QMSum | SAMSum | Trivia | PR-En | PR-Zh | Avg | ||
| FocusVTC w/o SFT (Ours) | Vision | 43.77 | 46.73 | 53.80 | 26.74 | 61.26 | 24.07 | 23.95 | 31.56 | 85.31 | 72.32 | 72.74 | 49.30 |
| FocusVTC w/o GRPO (Ours) | Vision | 34.52 | 39.94 | 31.77 | 14.73 | 56.54 | 17.47 | 16.69 | 31.64 | 89.23 | 67.12 | 16.85 | 37.86 |
| FocusVTC w/o GRPO+tools (Ours) | Vision | 16.46 | 32.70 | 23.16 | 8.42 | 49.70 | 2.96 | 3.79 | 3.88 | 56.42 | 65.19 | 16.83 | 25.41 |
| FocusVTC-50 (Ours) | Vision | 47.61 | 41.73 | 49.21 | 29.80 | 52.73 | 21.47 | 12.15 | 30.78 | 87.21 | 94.05 | 93.66 | 50.95 |
| FocusVTC-100 (Ours) | Vision | 46.90 | 45.06 | 54.22 | 30.10 | 61.14 | 22.47 | 20.06 | 32.76 | 90.10 | 95.73 | 96.00 | 54.05 |
| Model / stage | Input | Retrieval | Reasoning | Memory | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 8 | 16 | 32 | 64 | Avg | 8 | 16 | 32 | 64 | Avg | 8 | 16 | 32 | 64 | Avg | ||
| FocusVTC w/o SFT (Ours) | Vision | 99.83 | 85.55 | 82.22 | 81.07 | 87.17 | 50.88 | 49.09 | 33.96 | 18.65 | 38.15 | 31.24 | 22.21 | 21.57 | 19.20 | 23.56 |
| FocusVTC w/o GRPO (Ours) | Vision | 87.78 | 79.56 | 80.21 | 73.52 | 80.27 | 40.17 | 25.68 | 12.04 | 3.57 | 20.37 | 12.12 | 17.98 | 20.60 | 20.00 | 17.68 |
| FocusVTC w/o GRPO+tools (Ours) | Vision | 66.28 | 59.67 | 51.72 | 39.34 | 54.25 | 7.82 | 1.35 | 0.93 | 0.00 | 2.53 | 7.50 | 4.23 | 6.18 | 9.79 | 6.93 |
| FocusVTC-50 (Ours) | Vision | 96.96 | 80.65 | 81.97 | 65.56 | 81.29 | 38.51 | 29.37 | 14.22 | 7.00 | 22.28 | 17.63 | 13.65 | 18.06 | 24.99 | 18.58 |
| FocusVTC-100 (Ours) | Vision | 97.94 | 87.85 | 85.34 | 77.05 | 87.05 | 39.53 | 29.73 | 16.67 | 11.21 | 24.29 | 25.53 | 17.84 | 20.85 | 25.64 | 22.47 |
| Model / stage | Input | 2 needles | 4 needles | 8 needles | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 8 | 16 | 32 | 64 | 128 | 256 | Avg | 8 | 16 | 32 | 64 | 128 | 256 | Avg | 8 | 16 | 32 | 64 | 128 | 256 | Avg | ||
| LLaMA-3.1-8B ( Grattafiori et al., 2024 ) | Text | 54.27 | 53.21 | 51.05 | 29.81 | 24.98 | 20.90 | 39.04 | 33.42 | 25.97 | 22.73 | 26.97 | 12.68 | 6.00 | 21.30 | 23.80 | 17.69 | 19.85 | 17.72 | 11.79 | 7.80 | 16.44 |
| Qwen2.5-7B ( Yang et al., 2025b ) | Text | 45.92 | 51.07 | 46.97 | 34.67 | 37.57 | 37.60 | 42.30 | 25.96 | 20.13 | 19.93 | 24.25 | 17.29 | 12.30 | 19.98 | 17.64 | 19.48 | 12.41 | 14.80 | 14.24 | 13.70 | 15.38 |
| Qwen3-8B ( Yang et al., 2025a ) | Text | 58.95 | 41.18 | 36.18 | 24.99 | 20.89 | 17.50 | 33.28 | 29.34 | 22.67 | 20.34 | 23.63 | 19.11 | 15.50 | 21.77 | 18.75 | 19.69 | 16.81 | 17.86 | 15.00 | 12.60 | 16.79 |
| GLM-4-9B ( Team GLM et al., 2024 ) | Text | 39.77 | 15.87 | 18.42 | 18.63 | 18.42 | 18.20 | 21.55 | 15.17 | 13.78 | 9.18 | 20.27 | 15.05 | 11.20 | 14.11 | 14.55 | 9.65 | 9.34 | 9.47 | 8.97 | 8.50 | 10.08 |
| Qwen3.5-9B ( Qwen Team, 2026 ) | Text | 54.96 | 51.59 | 48.59 | 43.44 | 41.48 | 39.08 | 46.52 | 36.72 | 35.02 | 34.55 | 25.44 | 25.82 | 21.66 | 29.87 | 23.04 | 20.50 | 21.53 | 12.68 | 21.13 | 12.48 | 18.56 |
| Model | Single-1 | Single-2 | Single-3 | MKey-1 | MKey-2 | MKey-3 | MValue | MQuery | QA-1 | QA-2 | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3.5-9B | 54.00 | 48.00 | 0.00 | 43.00 | 40.00 | 1.00 | 37.25 | 42.75 | 62.00 | 46.00 | 37.40 |
| GLM-4.1V-9B | 61.00 | 26.00 | 0.00 | 30.00 | 19.00 | 0.00 | 19.50 | 33.25 | 50.00 | 48.00 | 28.68 |
| Glyph | 74.00 | 76.00 | 44.00 | 69.00 | 37.00 | 2.00 | 70.75 | 74.50 | 64.00 | 64.00 | 57.53 |
| FocusVTC w/o GRPO (Ours) | 50.00 | 39.00 | 0.00 | 29.00 | 37.00 | 0.00 | 24.25 | 24.50 | 65.00 | 54.00 | 32.28 |
| FocusVTC w/o GRPO+tools (Ours) | 43.00 | 29.00 | 1.00 | 24.00 | 37.00 | 1.00 | 23.75 | 23.00 | 22.00 | 26.00 | 22.98 |
| FocusVTC-50 (Ours) | 82.99 | 82.73 | 46.03 | 73.50 | 66.15 | 37.59 | 54.36 | 71.10 | 77.27 | 68.58 | 66.03 |
| Model | MK-NIAH | MV-NIAH | QA | Overall | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Basic | Easy | Medium | Hard | Basic | Easy | Medium | Hard | Basic | Easy | Medium | Hard | Avg | |
| Qwen3.5-9B | 19.00 | 79.44 | 70.00 | 46.00 | 10.75 | 8.65 | 23.36 | 23.45 | 69.00 | 34.00 | 71.00 | 81.00 | 44.64 |
| GLM-4.1V-9B | 8.00 | 51.12 | 29.00 | 25.00 | 3.50 | 9.67 | 10.64 | 18.24 | 53.00 | 32.00 | 53.58 | 53.50 | 28.94 |
| Glyph | 16.00 | 67.15 | 45.00 | 43.00 | 9.25 | 5.72 | 17.02 | 34.05 | 86.00 | 93.00 | 78.35 | 80.36 | 47.91 |
| FocusVTC w/o GRPO (Ours) | 18.00 | 59.96 | 51.00 | 39.00 | 7.25 | 7.05 | 16.79 | 27.24 | 74.00 | 83.00 | 68.50 | 71.17 | 43.58 |
| FocusVTC w/o GRPO+tools (Ours) | 20.00 | 71.32 | 45.00 | 33.00 | 10.75 | 7.03 | 13.04 | 29.78 | 72.00 | 36.00 | 63.27 | 71.68 | 39.41 |
| Model | OCRBench | DocVQA | MMMU | MME | ChartQA | InfoVQA |
|---|---|---|---|---|---|---|
| Qwen3.5-9B ( Qwen Team, 2026 ) | 851 | 92.38 | 65.12 | 2424.02 | 85.96 | 74.76 |
| Glyph ( Cheng et al., 2026 ) | 799 | 91.75 | 57.67 | 2253.13 | 72.76 | 71.89 |
| GLM-4.1V-9B ( V Team et al., 2025 ) | 820 | 91.34 | 64.33 | 2392.53 | 70.76 | 70.52 |
| FocusVTC (Ours) | 860 | 92.43 | 66.73 | 2457.62 | 86.28 | 76.20 |
| Model | VRAG | VH | NIAH image | NIAH text | Summ | DocQA |
| Qwen3.5-9B ( Qwen Team, 2026 ) | 61.12 | 56.10 | 45.04 | 71.82 | 29.19 | 68.96 |