SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models
Authors: Ahmadreza Jeddi, Enming Zhang, Jasper Gerigk, Hakki Karaimer, Mozhgan Nasr Azadani, Jiayun Luo, Minh Ngoc Le, Gholamali Aminian, +7 more
Organizations: University of Toronto · Vector Institute · Samsung AI Center Toronto · Stanford University · University of Waterloo · University of British Columbia · Alan Turing Institute · Tsinghua University · York University · NVIDIA
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, repeated sampling from the same pruned visual representation recovers many examples missed by greedy decoding, indicating that useful visual evidence can remain accessible but be used unreliably. We call this the representation-utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes. SCOPD requires no ground-truth responses, architectural changes, or additional inference-time computation. We further introduce SCOPD+, which uses a small visual-budget intervention to identify visually sensitive response positions and selectively distill them. At 10% visual-token retention, the Vanilla model retains 86.37% of its unpruned performance across 13 benchmarks. SCOPD raises this to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.
Figures & tables
Figure 1: Visual-token pruning creates a representation–utilization gap that sparse-context adaptation can recover. (a) Repeated sampling from the same sparse visual context can recover valid reasoning trajectories missed by greedy decoding. (b) Across 500 examples solved greedily by the unpruned model, Pass@ K substantially improves under aggressive VisionZip Yang et al. (2025) pruning, while a language-only control remains near zero, showing that useful visual evidence remains accessible. (c) SCOPD+ complements token pruning by adapting the model to better use this remaining evidence, substantially improving performance at aggressive token budgets.
Figure 2: Motivation and pipeline of SCOPD+ . (a) Large teacher–student KL can arise from language-level disagreement even when a prediction is weakly affected by visual context. We therefore use a small visual-budget intervention as a direct test of visual dependence: positions whose predictions change with additional visual evidence are more visually sensitive. (b) The student generates an on-policy trajectory at budget b . The same prefixes are scored by the student at budgets b and b+ and by the full-context teacher. We compute visual sensitivity using the Jensen–Shannon divergence between the two student distributions, select the top ρ=10% response positions, and apply the teacher-to-student KL loss only at those positions.
Table 3Table 4
Figure 3: Sensitivity of SCOPD+ to its hyperparameters. Results are normalized over 13 image benchmarks. (a) Performance peaks at ρ=10% , with no benefit from denser supervision. (b) Performance is stable across visual-budget interventions δ , with δ=1% performing best overall.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Recoverability under aggressive visual-token pruning. We report greedy and Pass@ K performance ( K∈{4,8,16,32,64} ) across the six sources in our 500-example diagnostic set. VisionZip retains 10% of visual tokens, while the language-only control receives no visual input. Performance consistently improves with sampling under the fixed sparse context, whereas the language-only control remains near zero.
Figure 5: Examples of reasoning-aware validity judgments. Four sampled trajectories illustrate cases where answer correctness alone can miss unsupported or visually inconsistent reasoning, motivating our GPT-based validity judge.
Figure 6: High teacher–student KL does not always imply visual relevance. Language-level differences can produce large KL despite low visual sensitivity.
Figure 7: Training dynamics of the GRPO baseline . Each point is a 128-step window; faint lines are 32-step windows. Accuracy is exact match on the prompts with a checkable answer (multiple choice, numeric, yes/no; 65% of prompts). (a) Mean accuracy and Pass@4 are flat at ≈ 76% and ≈ 91%. (b) Only ≈ 35% of groups contain both correct and incorrect responses and thus receive a non-zero advantage; ≈ 55% are already all correct. (c) Mean length stays at ≈ 160 tokens (shaded: interquartile range); incorrect responses are consistently longer than correct ones. (d) Format compliance improves slightly (84.8% → 89.2%); truncation stays below 3%.
Method
TFLOPs / Sample
Time / Sample
Peak Memory
SCOPD
54.25
17.21 s
25.02 GiB
SCOPD+
65.84
17.54 s
25.13 GiB
Appendix
Table 9: Training overhead and response length. Compute measurements use batch size 32 on a single NVIDIA L40 GPU. TFLOPs and time are reported per sample; peak memory is measured over the full batch.
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.
Qingchan Zhu, Weihang You, Hanqi Jiang +3
School of Computing, University of Georgia · College of Engineering, Northeastern University
Are low-attention visual tokens truly redundant in vision-language reasoning? Existing pruning methods often assume so, ranking visual tokens by shallow text-to-image attention and discarding low-scoring patches to accelerate LVLM inference. We show that this scalar criterion is unreliable for compositional reasoning: tokens ignored in early layers can later become essential for resolving secondary objects, spatial relations, and contextual cues. Premature pruning can therefore induce Visual Aphasia, a failure mode in which the model loses visual grounding and falls back on language priors. We introduce COAST (COntrastive Adaptive Semantic Token Pruning), a training-free pruning framework that casts compression as adaptive semantic routing. COAST uses native cross-modal attention to identify query-specific anchors and estimate contextual dispersion via attention entropy, then adapts the retention trade-off between semantic evidence and spatial context. It further uses a contrastive routing score to preserve both anchor-aligned evidence and complementary spatial context. Across seven benchmarks, COAST reduces visual tokens by 77.8% and achieves a 2.15x latency speedup while retaining 98.64% of the original average performance. Beyond a single backbone or compression setting, COAST consistently outperforms strong pruning baselines across token budgets and generalizes across multiple LVLM families, showing that adaptive semantic routing is a robust alternative to one-shot scalar pruning
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.