UniCounting: Instance-Aware Proposal Consolidation for Image-Query-Free Multi-Category Counting
Organizations: Shenzhen University · The Hong Kong University of Science and Technology (Guangzhou) · Hunan University of Science and Technology · Peng Cheng Laboratory · Fudan University
Abstract
Visual counting is commonly formulated as counting a single specified target, with a model receiving an image-specific exemplar, text query, or target category and returning a single count. We instead study fixed-vocabulary image-query-free multi-category counting. A global vocabulary is fixed for each run, and, given only an RGB image, the model predicts a complete category--count vector without being told which categories appear. We present UniCounting, which casts counting as instance-aware structural inference over an over-complete proposal set. Generic segmenters produce duplicate masks, partial views, and proposals from neighboring instances; semantic scores can name them but cannot determine which denote the same object. Frozen SAM~2.1 generates masks, while frozen DINOv2 and OpenCLIP provide relation and category features. A 3,267-parameter category-shared relation head predicts same-instance affinities from instance-mask-derived supervision. Sparse graph construction, representative selection, labeling, and background-margin admission then convert each admitted component into one count with replayable group evidence. Only the relation head is trained, without count or density-map targets. On COCO clean500, UniCounting obtains lower point-estimate vector error and absent-class false mass than calibrated OWLv2-All80, with comparable micro presence F1. Under a matched decoder, the learned relation reduces both errors relative to mask containment, mask IoU, CLIP, and DINO, while revealing a fragmentation--merge trade-off. We also report transfer diagnostics on OmniCount-sub, FSC-147, and CARPK.
Figures & tables
| Method | Vector | Present MAE | False mass | Total MAE | Macro-F1 | Micro-F1 |
|---|---|---|---|---|---|---|
| Zero vector | 7.018 | 2.425 | 0.000 | 7.018 | 0.000 | 0.000 |
| Fixed-list detector (OWLv2-All80) | 6.860 | 2.205 | 0.474 | 5.800 | 0.472 | 0.415 |
| UniCounting | 6.485 0.058 | 2.129 0.005 | 0.325 0.062 | 5.868 0.050 | 0.407 0.007 | 0.416 0.005 |
| Faster R-CNN (fully supervised ref.) | 3.098 | 0.916 | 0.446 | 2.078 | 0.809 | 0.832 |
| Edge/grouping rule | Vector | False mass | Total MAE | Micro-F1 | |
|---|---|---|---|---|---|
| No grouping / singletons | – | 29.187 | 20.347 | 24.440 | 0.319 |
| Mask containment | 0.005 | 9.935 | 4.697 | 5.020 | 0.427 |
| Mask IoU | 0.005 | 13.649 | 8.265 | 7.965 | 0.374 |
| CLIP cosine | 0.500 | 7.335 | 1.452 | 4.987 | 0.431 |
| DINO cosine | 0.310 | 6.814 | 0.809 | 5.463 | 0.434 |
| Relation-head score overlap guard | 0.465 | 9.597 | 4.361 | 4.635 | 0.414 |
| Stage | Selected variant | Vector | Micro-F1 |
|---|---|---|---|
| Graph | same-instance edges | ||
| Represent. | max category margin | ||
| Labeling | raw cosines | ||
| Admission | background-margin |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| SAM 2.1 ( Ravi et al. 2025 ) checkpoint | sam2.1-hiera-small , SHA-256 prefix 0a4067b1 |
| Input resolution | checkpoint-native preprocessing |
| Points per side / crop layers | 24 / 0 |
| Predicted-IoU threshold | 0.7 |
| Stability threshold | 0.8 |
| Box NMS threshold | 0.7 |
| Setting | Value |
|---|---|
| GPU / memory | NVIDIA GeForce RTX 4090 / 24,564 MiB (approximately 24 GB) |
| CPU | Intel Core i9-14900KF |
| System memory | 62 GiB available (approximately 64 GB physical memory) |
| Operating system | Ubuntu 24.04.4 LTS under WSL2 |
| Linux kernel | 6.6.114.1- microsoft-standard-WSL2 |
| Python | 3.12.3 |
| Edge score | threshold | vector | micro-F1 |
|---|---|---|---|
| Mask containment | 0.005 | 12.5333 | 0.4289 |
| Mask IoU | 0.005 | 15.8222 | 0.3920 |
| CLIP cosine | 0.500 | 7.2444 | 0.4632 |
| DINO cosine | 0.310 | 7.1185 | 0.4160 |
| Relation-head score | 0.465 | 6.7930 | 0.4260 |
| Method | Vector | Micro-F1 |
|---|---|---|
| Zero vector | 6.447 | 0.000 |
| UniCounting | 6.715 0.123 | 0.283 0.011 |
| Method / seed | TP | FP | FN | Micro-P | Micro-R | Micro-F1 |
|---|---|---|---|---|---|---|
| OWLv2-All80 | 432 | 202 | 1015 | 0.681 | 0.299 | 0.415 |
| UniCounting / 17 | 429 | 187 | 1,018 | 0.696 | 0.296 | 0.416 |
| UniCounting / 42 | 418 | 165 | 1,029 | 0.717 | 0.289 | 0.412 |
| UniCounting / 73 | 420 | 125 | 1,027 | 0.771 | 0.290 | 0.422 |
| Grouping rule | Pair F1 | Purity | Merge error | Comp./instance | Fragmented | Rep. IoU |
|---|---|---|---|---|---|---|
| No grouping / singletons | 0.0000 | 1.0000 | 0.0000 | 4.0770 | 0.6320 | 0.1680 |
| Mask containment | 0.6910 | 0.9090 | 0.1150 | 1.4150 | 0.1350 | 0.2020 |
| Mask IoU | 0.4984 | 0.9558 | 0.0515 | 1.8104 | 0.2282 | 0.1758 |
| CLIP cosine | 0.5040 | 0.6740 | 0.4680 | 1.6320 | 0.4300 | 0.1450 |
| DINO cosine | 0.6270 | 0.6510 | 0.5060 | 1.1950 | 0.1690 | 0.1760 |
| Relation-head score overlap guard | 0.6820 | 0.8500 | 0.1660 | 1.4180 | 0.1830 | 0.2030 |
| Inner45 | COCO clean500 | OmniCount-sub 2K | |||||
|---|---|---|---|---|---|---|---|
| Stage | Variant | Micro-F1 | Micro-F1 | Micro-F1 | |||
| Graph | hard semantic instance | 32.978 0.868 | 0.300 0.003 | 30.597 1.040 | 0.291 0.003 | 26.553 1.073 | 0.144 0.003 |
| same-instance edges, all | 6.889 0.000 | 0.365 0.000 | 6.686 0.000 | 0.342 0.000 | 6.769 0.000 | 0.257 0.000 | |
| same-instance edges, top2 | 6.867 0.059 | 0.417 0.005 | 6.810 0.010 | 0.350 0.002 | 6.976 0.009 | 0.250 0.002 | |
| Represent. | maximum area | 7.296 0.078 | 0.332 0.008 | 6.837 0.011 | 0.345 0.002 | 7.293 0.012 | 0.144 0.003 |
| learned completeness | 7.400 0.102 | 0.323 0.007 | 6.858 0.003 | 0.339 0.003 | 7.305 0.014 | 0.141 0.003 | |