cs.CVOct 7, 2026

InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation

Authors: Yuchen Li, Shaoyang Zhou, Yiran Wang, Ruiyi Deng, Haoyu Wang, Ziru Wei, Zhen Zhao, Luping Zhou

Organizations: The University of Sydney · The ATLAS Institute · Shanghai Artificial Intelligence Laboratory

Abstract

Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors. We introduce InstanceBench, an instance-centered diagnostic benchmark comprising 6,194 images, 9,264 target instances, and 25,077 human-verified expressions. Each target-centric expression set (TCES) fixes the image and target mask while pairing a minimal expression with a same-target variant that uses another valid cue or grounding path. A compact referential-logic taxonomy spans direct target evidence, same-class selection, relational and compositional grounding, and exclusion, while logic-critical construction suppresses simpler shortcuts. Identity-aware metrics measure target retention and set-level success while separating selection from mask-generation errors. Across 22 native-mask RES checkpoints from 18 model families, the strongest checkpoint reaches 67.1% mIoU but only 59.6% [email protected]. Controlled interventions confirm language sensitivity, while failure decomposition identifies target selection rather than mask decoding as the main bottleneck. On a controlled training subset, matched supervision improves identity-aware performance, showing that the diagnosed capability responds to targeted supervision. Collectively, InstanceBench supports a measure-diagnose-improve cycle: measuring target consistency across grounding paths, localizing failure sources, and evaluating targeted interventions.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning

    Aug 25, 2026Shuaishuai Cao, Min Huang, Meng Tang +3Spatial Reasoning BenchmarksGeospatial Reasoning

  2. SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction

    May 19, 2026Zhixiong Zhang, Yizhuo Li, Shuangrui Ding +6Image SegmentationReferring Video Object Segmentation

  3. ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring

    May 8, 2026Tianhao Niu, Ziyu Han, Xuan Dong +2Referring Image SegmentationVisual Grounding