cs.CVSep 3, 2026

Efficient Semantic Understanding from Digital Foveation

Authors: Caterina CaccavellaVittorio FraAndreas ZieglerGiulia D'AngeloYulia Sandamirskaya

Organizations: Zurich University of Applied Sciences (ZHAW), Wädenswil, Switzerland · ETH Zürich, Zürich, Switzerland · Politecnico di Torino, Turin, Italy · University of Tübingen, Tübingen, Germany · Czech Technical University, Prague, Czech Republic

Abstract

Dense semantic segmentation allocates computational resources uniformly across the entire image, regardless of scene complexity or task relevance. Inspired by biological vision, we investigate whether semantic understanding can be achieved more efficiently through digital foveated perception. We introduce a lightweight active-vision pipeline that combines saliency-driven fixation selection, high-resolution foveal observations, low-resolution contextual information, semantic accumulation, and adaptive computation. Beyond conventional dense prediction metrics, we use object-level evaluation to measure semantic understanding under sparse observations. On ADE20K-Object, a single foveated observation achieves 95.9% of the baseline Top-1 accuracy and 96.9% of the baseline Top-3 accuracy while requiring only 4.7% of the computational cost. At the scene level, semantic accumulation recovers 90.6% of the baseline object recall while using 58.6% of the computation. These results suggest that substantial semantic understanding can emerge from sparse observations when computation is allocated selectively, highlighting active vision as an efficient alternative to uniform dense processing and motivating evaluation protocols beyond conventional pixel-wise segmentation metrics.

Explore similar work

CardsList
  1. Policy-based Foveated Imaging and Perception

    Jun 1, 2026Howard Xiao, Jan Ackermann, Boyang Deng +1Vision SensorsLow-Resolution