cs.CVOct 6, 2026

UniCounting: Instance-Aware Proposal Consolidation for Image-Query-Free Multi-Category Counting

Authors: Jinshi Liu, Pan Liu, Lei He, Weichao Luo, Rui Qian

Organizations: Shenzhen University · The Hong Kong University of Science and Technology (Guangzhou) · Hunan University of Science and Technology · Peng Cheng Laboratory · Fudan University

Abstract

Visual counting is commonly formulated as counting a single specified target, with a model receiving an image-specific exemplar, text query, or target category and returning a single count. We instead study fixed-vocabulary image-query-free multi-category counting. A global vocabulary is fixed for each run, and, given only an RGB image, the model predicts a complete category--count vector without being told which categories appear. We present UniCounting, which casts counting as instance-aware structural inference over an over-complete proposal set. Generic segmenters produce duplicate masks, partial views, and proposals from neighboring instances; semantic scores can name them but cannot determine which denote the same object. Frozen SAM~2.1 generates masks, while frozen DINOv2 and OpenCLIP provide relation and category features. A 3,267-parameter category-shared relation head predicts same-instance affinities from instance-mask-derived supervision. Sparse graph construction, representative selection, labeling, and background-margin admission then convert each admitted component into one count with replayable group evidence. Only the relation head is trained, without count or density-map targets. On COCO clean500, UniCounting obtains lower point-estimate vector ℓ1\ell_1 error and absent-class false mass than calibrated OWLv2-All80, with comparable micro presence F1. Under a matched decoder, the learned relation reduces both errors relative to mask containment, mask IoU, CLIP, and DINO, while revealing a fragmentation--merge trade-off. We also report transfer diagnostics on OmniCount-sub, FSC-147, and CARPK.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Count Anything

    May 29, 2026Mengqi Lei, Shuokun Cheng, Wei Bao +4CountingVision Datasets

  2. Spatially-Aware Class-Agnostic Object Counting

    Jul 18, 2026Robert Wijaya, Md. Tanvir Hossain, Amanda Kau +1Feature PyramidCounting

  3. Count Anything at Any Granularity

    May 11, 2026Chang Liu, Haoning Wu, Weidi XieCountingFine-Grained Perception