InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation
Organizations: The University of Sydney · The ATLAS Institute · Shanghai Artificial Intelligence Laboratory
Abstract
Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors. We introduce InstanceBench, an instance-centered diagnostic benchmark comprising 6,194 images, 9,264 target instances, and 25,077 human-verified expressions. Each target-centric expression set (TCES) fixes the image and target mask while pairing a minimal expression with a same-target variant that uses another valid cue or grounding path. A compact referential-logic taxonomy spans direct target evidence, same-class selection, relational and compositional grounding, and exclusion, while logic-critical construction suppresses simpler shortcuts. Identity-aware metrics measure target retention and set-level success while separating selection from mask-generation errors. Across 22 native-mask RES checkpoints from 18 model families, the strongest checkpoint reaches 67.1% mIoU but only 59.6% [email protected]. Controlled interventions confirm language sensitivity, while failure decomposition identifies target selection rather than mask decoding as the main bottleneck. On a controlled training subset, matched supervision improves identity-aware performance, showing that the diagnosed capability responds to targeted supervision. Collectively, InstanceBench supports a measure-diagnose-improve cycle: measuring target consistency across grounding paths, localizing failure sources, and evaluating targeted interventions.
Figures & tables
| Data | Diagnostic construction | Diagnostic evaluation | ||||||
| Benchmark | Real images | Pixel masks | Taxonomy scale | Controlled construction | Controlled same-target grounding-path variants | Logic-wise diagnosis | Set-level reliability | Selection–mask decomposition |
| RefCOCO/+/g | – | |||||||
| CLEVR-Ref+ | 5 categories | |||||||
| Cops-Ref | 6 logic forms | |||||||
| FineCops-Ref | 6 structures | |||||||
| gRefCOCO | – | |||||||
| Model / checkpoint | Venue | Group | oIoU | mIoU | [email protected] | [email protected] | [email protected] | Worst-path IoU |
|---|---|---|---|---|---|---|---|---|
| SAMTok-CO | CVPR’26 | U | 57.56 | 67.14 | 71.18 | 59.60 | 32.42 | 56.28 |
| X2SAM | ECCV’26 | U | 55.77 | 65.92 | 68.50 | 56.60 | 38.73 | 54.43 |
| UniPixel-7B | NeurIPS’25 | U | 54.68 | 62.78 | 63.18 | 51.82 | 24.61 | 52.07 |
| Sa2VA-2B | TPAMI’26 | U | 53.44 | 58.83 | 59.13 | 45.37 | 21.83 | 45.87 |
| X-SAM | AAAI’26 | U | 39.91 | 56.81 | 58.16 | 46.60 | 30.27 | 45.28 |
| InstructSeg | ICCV’25 | U | 41.55 | 52.73 | 53.51 | 39.07 | 21.38 | 38.52 |
| Model | Para. Ret. | Variant Ret. | Para. NCI | Variant NCI | NCI [95% CI] | All-target-ID | Cond. mask IoU |
|---|---|---|---|---|---|---|---|
| SAMTok | 97.55 | 88.47 | 0.565 | 0.151 | [ ] | 62.47 | 97.48 |
| X2SAM | 96.75 | 89.18 | 0.662 | 0.201 | [ ] | 58.32 | 98.05 |
| UniPixel | 97.42 | 89.16 | 0.770 | 0.287 | [ ] | 58.98 | 92.31 |
| Sa2VA | 96.17 | 85.43 | 0.593 | 0.147 | [ ] | 51.48 | 90.16 |
| ConverSeg | 96.11 | 87.22 | 0.660 | 0.340 | [ ] | 51.62 | 90.30 |
| PaDT | 96.07 | 83.84 | 0.644 | 0.185 | [ ] | 50.12 | 94.56 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Subtype | Example | Identification requirement |
| Same-class set selection | Ordinal | “the second wine bottle from the left” | Order the same-class instances and select a specified index. |
| Extremal | “the leftmost visible person” | Select the instance attaining a positional or visual extremum. | |
| Relative selection | “the larger of the two pizza slices” | Compare same-class instances using a visible property or distance. | |
| Subgroup-based | “the right pale-pink cruller among the two at the bottom” | Form a constrained subgroup before selecting one member from it. | |
| Single-anchor relation | – | “the young girl immediately to the left of the woman cutting the pizza” | Resolve one independently identifiable anchor and apply one spatial, distance, contact, containment, facing, looking, or body-orientation relation. |
| Multi-anchor compositional grounding | Multi-anchor relation | “the dark suitcase between the wooden crate and the tan suitcase” | Jointly resolve at least two anchors whose combined constraints identify the target. |
| Model | Group | Language backbone | Extra | Audited downstream training scope |
|---|---|---|---|---|
| X2SAM | U | Qwen3-VL | Yes | RefCOCO/+/g plus generic, reasoning, grounded-conversation, interactive, and visual-grounded segmentation over images and videos. |
| SAMTok | U | MLLM | Yes | Large-scale mask-tokenizer pretraining and mask–text conversations spanning RES, instance/scene parsing, region QA, and interactive reasoning. |
| UniPixel | U | MLLM | Yes | Joint image/video referring, reasoning, video segmentation, captioning, and pixel-QA-related tasks. |
| Sa2VA | U | Qwen3-VL | Yes | Image/video instruction tuning for chat, RES, referring-video segmentation, and grounded caption generation. |
| X-SAM | U | MLLM | Yes | Conversation alignment followed by generic, referring, reasoning, grounded-conversation, interactive, and visual-grounded segmentation. |
| HyperSeg / InstructSeg | U | MLLM | Yes | Joint generic/referring/reasoning and image/video segmentation mixtures. |
| Localization and mask quality | TCES reliability | |||||
|---|---|---|---|---|---|---|
| Model | Box [email protected] | oIoU | mIoU | [email protected] | All-target-ID | Variant ret. |
| Qwen2.5-VL-7B | 66.91 | 43.45 | 53.20 | 42.21 | 52.27 | 77.59 |
| Qwen2.5-VL-32B-AWQ | 73.57 | 49.20 | 58.98 | 48.99 | 59.73 | 81.42 |
| Qwen2.5-VL-72B-AWQ | 75.64 | 52.07 | 61.01 | 51.46 | 61.44 | 81.73 |
| Qwen3-VL-4B | 73.83 | 53.90 | 63.20 | 53.99 | 59.94 | 78.85 |
| Qwen3-VL-8B-Thinking | 71.36 | 55.56 | 60.84 | 52.04 | 57.67 | 79.14 |
| Model | Expr.-ID | Minimal-ID | Variant-ID | All-target-ID | Variant ret. |
|---|---|---|---|---|---|
| Qwen2.5-VL-7B | 43.95 | 46.69 | 41.21 | 30.61 | 65.57 |
| Qwen2.5-VL-32B-AWQ | 53.38 | 55.09 | 51.67 | 38.91 | 70.63 |
| Qwen2.5-VL-72B-AWQ | 54.11 | 56.03 | 52.20 | 41.63 | 74.30 |
| Qwen3-VL-8B-Thinking | 45.08 | 50.24 | 39.92 | 29.50 | 58.71 |
| Qwen3.8-27B | 66.07 | 68.41 | 63.74 | 54.22 | 79.26 |
| Gemini-3.8-Flash | 76.29 | 77.30 | 75.28 | 65.17 | 84.30 |
| Architecture | Supervision | mIoU | [email protected] | All-target-ID |
|---|---|---|---|---|
| LAVT | Source text | |||
| 50/50 mixed | ||||
| Logic-aware tracks | ||||
| DETRIS | Source text | |||
| 50/50 mixed | ||||
| Logic-aware tracks |
| Model | Target evidence (1,273) | Single anchor (1,830) | Extremal (401) | Subgroup (321) | Multi- anchor (635) | Ordinal (403) | Exclusion (653) | Nested (207) | Relative (113) |
|---|---|---|---|---|---|---|---|---|---|
| SAMTok | 78.84 | 67.65 | 76.81 | 65.50 | 60.32 | 57.82 | 54.56 | 52.14 | 69.23 |
| X2SAM | 79.19 | 66.00 | 73.75 | 60.59 | 58.66 | 57.92 | 53.95 | 47.84 | 74.07 |
| UniPixel | 75.41 | 64.04 | 73.30 | 57.23 | 57.09 | 45.59 | 51.77 | 46.73 | 64.66 |
| X-SAM | 72.69 | 57.06 | 71.45 | 49.43 | 47.47 | 48.07 | 40.95 | 33.25 | 61.12 |
| Sa2VA | 76.21 | 58.04 | 71.67 | 52.89 | 49.31 | 42.87 | 45.37 | 42.91 | 64.34 |
| ConverSeg | 70.89 | 57.35 | 67.51 | 51.23 | 47.34 | 43.54 | 46.17 | 40.83 | 60.16 |
| Treated control | X2SAM | SAMTok | UniPixel | X-SAM | |
|---|---|---|---|---|---|
| Ordinal extremal | 349 | ||||
| Multi-anchor spatial | 618 | ||||
| Nested spatial | 200 |
| Category only | Text swap | Inference-time aid ( mIoU) | |||||
|---|---|---|---|---|---|---|---|
| Model | mIoU | Drop | Donor IoU | Donor pref. | Logic-explicit | Deliberation | SC context |
| X2SAM | 24.29 | 80.61 | 93.50 | ||||
| SAMTok | 24.84 | 79.46 | 94.23 | ||||
| UniPixel | 21.24 | 77.12 | 93.07 | ||||
| X-SAM | 22.34 | 76.60 | 90.29 | ||||
| Model / checkpoint | Minimal | Paraphrase | Same-target | Exclusion | Paired exclusion drop |
|---|---|---|---|---|---|
| SAMTok-CO | 69.73 | 82.94 | 66.36 | 57.76 | 13.03 |
| X2SAM | 68.74 | 82.22 | 65.08 | 56.90 | 13.81 |
| UniPixel-7B | 65.00 | 80.00 | 62.25 | 54.46 | 13.96 |
| Sa2VA-2B | 62.22 | 78.63 | 56.97 | 48.04 | 17.52 |
| X-SAM | 59.30 | 79.59 | 56.57 | 44.02 | 14.42 |
| ConverSeg-3B | 59.06 | 75.59 | 56.35 | 48.35 | 12.43 |
| Expression-level mIoU | TCES-level [email protected] | |||||
|---|---|---|---|---|---|---|
| Model | COCO | OI | COCO | OI | ||
| SAMTok-CO | 76.86 | 57.35 | 75.34 | 43.74 | ||
| X2SAM | 77.80 | 53.96 | 75.46 | 37.62 | ||
| UniPixel-7B | 74.61 | 50.86 | 70.15 | 33.36 | ||
| Sa2VA-2B | 71.82 | 45.74 | 64.69 | 25.93 | ||
| ConverSeg-3B | 70.07 | 43.52 | 61.95 | 24.00 | ||
| Model | Para. Ret. | Variant Ret. | Excl. Ret. | Para. [email protected] | Variant [email protected] | Excl. [email protected] |
|---|---|---|---|---|---|---|
| SAMTok | 97.45 | 83.51 | 75.57 | 97.12 | 83.64 | 75.17 |
| X2SAM | 96.67 | 81.81 | 71.03 | 96.43 | 81.96 | 72.34 |
| UniPixel | 97.44 | 83.06 | 73.05 | 96.50 | 81.66 | 73.01 |
| Sa2VA | 95.97 | 77.73 | 67.03 | 95.34 | 74.99 | 63.75 |
| ConverSeg | 96.02 | 79.01 | 69.41 | 95.10 | 76.58 | 64.03 |
| PaDT | 96.09 | 78.60 | 68.77 | 95.98 | 82.57 | 71.77 |
| Ordinal extremal | Nested single-anchor | ||||
|---|---|---|---|---|---|
| Model | COCO | OI | COCO | OI | |
| SAMTok | 0.867 | ||||
| X2SAM | 0.833 | ||||
| UniPixel | 0.950 | ||||
| X-SAM | 0.750 | ||||
| Sa2VA | 0.833 | ||||
| Source | Logic category | Identity error | Same-class switch | Correct ID, poor mask | |
|---|---|---|---|---|---|
| COCO | Single-anchor | 984 | 66.46 | 59.45 | 33.54 |
| COCO | Multi-anchor | 205 | 77.56 | 69.76 | 22.44 |
| COCO | Nested | 132 | 78.79 | 70.45 | 21.21 |
| OpenImages | Single-anchor | 2,289 | 76.45 | 51.99 | 23.55 |
| OpenImages | Multi-anchor | 1,211 | 81.34 | 56.48 | 18.66 |
| OpenImages | Nested | 468 | 73.50 | 47.65 | 26.50 |
| Logic category | Identity error | Same-class switch | Correct ID, poor mask |
|---|---|---|---|
| Single-anchor | 73.45 | 54.23 | 26.55 |
| Multi-anchor | 80.79 | 58.40 | 19.21 |
| Nested | 74.67 | 52.67 | 25.33 |
| Min. score | Min. margin | Core-pair All- target-ID | Identity error | Same-class switch | Unresolved | Correct ID, poor mask |
|---|---|---|---|---|---|---|
| 0.05 | 0.00 | 59.77 | 73.28 | 58.04 | 15.24 | 26.72 |
| 0.05 | 0.05 | 59.04 | 75.31 | 55.51 | 19.80 | 24.69 |
| 0.10 | 0.00 | 59.60 | 73.64 | 57.67 | 15.98 | 26.36 |
| 0.10 | 0.05 | 58.92 | 75.55 | 55.17 | 20.38 | 24.45 |
| 0.10 | 0.10 | 58.53 | 76.73 | 53.73 | 22.99 | 23.27 |
| 0.15 | 0.05 | 58.86 | 75.80 | 54.81 | 20.99 | 24.20 |