cs.CVDate pending

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

Authors: Anindya MondalSauradip NagAnjan Dutta

Abstract

We present ABACUS, a unified vision-language model that jointly addresses object counting, crowd counting, referring-expression counting, and count-faithful image generation within a single 3B-parameter model. ABACUS introduces three contributions: density-aware adaptive zooming paired with an objectness map from multi-head self-attention decomposition to spatially ground count predictions; a boundary-aware count policy trained via GRPO with nested local, boundary, and global rewards to eliminate over- and undercounting at crop boundaries; and a cycle-consistent GRPO strategy in which the frozen understanding branch scores generated candidates on count-deviation and aesthetic quality, closing the understanding-generation synergy gap without any external critic or annotation. ABACUS achieves state-of-the-art results across seven benchmarks spanning object counting (FSC-147, CARPK), crowd counting (ShanghaiTech A/B), referring-expression counting (REC-8K), count-faithful generation (CoCoCount, T2I-CompBench, GenEval), and count reasoning (CountQA), surpassing both task-specific specialists and larger generalist models. Project page is at https://mondalanindya.github.io/pages/ABACUS.

Explore similar work

CardsList
  1. Count Anything

    May 29, 2026Mengqi Lei, Shuokun Cheng, Wei Bao +4CountingMicroscopy

  2. Vision as Unified Multimodal Generation

    Jul 7, 2026Xiaoyang Han, Jianhua Li, Kewang Deng +14Multimodal GenerationComputer Vision