cs.CVSep 24, 2026
SaveExploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge
Organizations: University of Illinois Urbana-Champaign
Abstract
Onboard vision-language models could enable satellites to answer queries directly, but exhaustive tiled inference over high-resolution imagery is slow and energy-intensive. We identify answer-invariant token redundancy (AITR): image tiles and vision tokens that can be removed without changing the final answer. We present Rift, a two-stage system that performs query-conditioned tile pruning followed by elastic prefill to reduce token budget. We evaluate it on LLaVA-1.5 7B running on Jetson AGX Orin. Compared with exhaustive tiled inference, Rift reduces energy by 78% and latency by 69%, while increasing accuracy from 45% to 73%.
Figures & tables
Figure 1 : Left: A tiled satellite scene and query; red boxes mark tiles irrelevant to the query. Right (bottom): Attention is concentrated on a small subset of the vision tokens within one tile. Right (top): Stage-level latency and energy, showing prefill dominates this input-heavy VLM inference.
Figure 2 : Rift reduces answer-invariant tile and token redundancy with its 2-stage framework: (i) Tile Pruning, and (ii) Elastic Prefill.
Figure 3 : ( Left ) Avg. per-query energy, latency and accuracy of Rift compared with a naive system and one that only prunes out redundant tiles. ( Right ) Impact of elastic prefill: Comparing avg. vlm inference energy, latency and accuracy with fixed configurations per-tile.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Step 1. Tile the region tiles of px
Figure 5 : Step 2. Per-tile relevance score
Figure 6 : Step 3. Prune irrelevant tiles
Figure 7 : Step 4. Elastic prefill decision policy picks per tile token budget
Figure 8 : Step 5. Per-tile VLM answers counts “3–9” (GT: 3–9)