cs.CVSep 28, 2026

DORA: Dynamic Online Reinforcement Agent for Token Pruning in Vision Transformers

Authors: Kaixuan He, Song Chen, Yi Kang

Organizations: University of Science and Technology of China, Hefei, China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, Hefei, China

Abstract

Vision Transformers (ViTs) incur quadratic self-attention cost in the number of tokens. Most token-reduction methods adapt token identities within a prescribed layer-wise compression schedule, or search a static mask offline, and thus limit online adaptation of when and how much to prune. We propose DORA (Dynamic Online Reinforcement Agent), which learns an input-adaptive pruning policy itself for frozen ViTs. At each eligible block, a hierarchical actor decides whether to prune, how many tokens to remove, and which tokens to remove from each image's evolving representation. Because early deletions change the states observed by later decisions, DORA formulates pruning as a finite-horizon Markov decision process. Complete-prefix shadow evaluations convert final-prediction fidelity into localized per-step credit, while closed-loop accuracy feedback adjusts the fidelity penalty toward a shared accuracy-drop target. A privileged critic and all shadow computations are training-only. Deployment retains the frozen backbone and a lightweight actor that applies hard deletion and packed variable-length FlashAttention, converting token reduction into measured speedups. On ImageNet-1K with DeiT-Base, DORA reduces FLOPs by 38.4% relative to the uncompressed backbone within one percentage point of accuracy loss. Averaged across four ViT-type backbones at matched accuracy, DORA uses 13.2% fewer FLOPs and achieves 32.4% higher throughput than the corresponding per-backbone baseline means. Under zero-shot transfer to ImageNet-A, these gains widen to 20.3% and 45.6%, respectively.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DORA: Dynamic Online Reinforcement Agent for Token Merging in Vision Transformers

    May 12, 2026Kaixuan He, Song Chen, Yi KangVision TransformerSoft-Token Representations

  2. ASAP: Attention Sink Anchored Pruning

    May 21, 2026Jaehyuk Lee, Hanyoung Kim, Yanggee Kim +1Visual Token PruningVision Transformer