We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer from fly and mouse to human and zebrafish. Most models were at chance zero-shot; a few examples helped mainly the closed and largest open ones. LoRA on a few thousand labels brought open models level with specialist models. When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors. The project will be publicly available upon acceptance.
Figures & tables
Figure 1: Synapse detection example inputs of fly. Left: presence , the query point marked by the yellow circle (answer: yes). Right: polarity , markers A and B inside the two partner processes (answer: B).
zero-shot
4-shot
LoRA
Model
fly P
fly Pol
mouse P
mouse Pol
fly P
fly Pol
mouse P
mouse Pol
fly P
fly Pol
mouse P
mouse Pol
Qwen2.5-VL-3B
49.1
50.0
49.5
50.0
50.9
50.0
53.9
50.0
50.3
57.8
74.4
76.4
Qwen2.5-VL-7B
50.0
49.0
49.6
49.2
48.8
49.1
48.2
49.3
71.8
70.0
86.3
81.4
Qwen2.5-VL-32B
48.2
49.1
46.2
49.8
55.2
49.4
52.2
50.3
68.2
61.7
83.5
77.4
Qwen3-VL-2B
51.2
49.1
48.3
50.9
50.0
48.7
54.1
50.0
69.7
56.9
81.0
79.9
Qwen3-VL-4B
53.3
48.2
47.5
51.9
59.1
51.7
48.8
49.8
72.4
56.3
83.7
82.1
Table 1: Synapse detection, balanced accuracy (%) on the test splits. P: presence ; Pol: polarity . Bold: best VLM per column; underlined : second best; Italic rows : specialist models, not ranked; –: not applicable.
Figure 2: Proofreading example inputs of fly. (a) A split-error site: the two segments are tinted red and blue; the proofreaders merged them, so the answer is yes. (b) A merge-error site: the segment is tinted red; the proofreaders split it, so the answer is no.
zero-shot
4-shot
LoRA
Model
S
M
S
M
S
M
Qwen3.5-9B
50.0
52.9
58.7
57.3
94.6
90.2
InternVL3-8B
51.0
50.0
49.0
57.3
93.9
90.1
Gemma 4 12B
50.1
50.0
62.7
58.0
95.0
90.9
CB2 ViT-B, EM only
–
–
–
–
94.2
91.1
Table 2: Proofreading on the complete ConnectomeBench2, balanced accuracy (%) on the test split: all 101,929 samples across 4 species. S: split-error subtask, M: merge-error subtask. Bold: best VLM per column; Italic row : the dataset’s own reported EM-only ViT-B; –: not applicable.
Qwen3.5-9B
InternVL3-8B
Gemma 4 12B
GPT-5.6 Sol
Claude Opus 5
ResNet-50
ConvNeXt-T
ViT-S/14
zero-shot
mouse S
50.0
53.4
50.4
67.2
73.1
–
–
–
mouse M
59.3
50.0
50.0
67.6
67.3
–
–
–
fly S
50.0
52.0
50.0
63.0
64.7
–
–
–
fly M
51.0
50.0
50.0
57.0
57.0
–
–
–
4-shot
Table 3: Proofreading on the subset of ConnectomeBench2 (Sec. 2.2 ), balanced accuracy (%). S: split-error subtask, M: merge-error subtask. Bold: best VLM per row; italic columns : specialist models, not ranked; –: not applicable.
Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both real-world knowledge bias and the scarcity of such images in training data. We present MissingBench-Verified, a benchmark designed to evaluate a specific and practically relevant scenario: when vision-language models fail to recognize that an essential component of an object has been removed. Across ten leading models, we observe consistent and significant failure rates that persist even when external tool evidence explicitly contradicts the model's visual perception. We further ask whether granting models access to image processing tools (e.g., cropping, contrast adjustment) enables autonomous inspection to resolve these failures. We find that existing mitigation strategies, including tool-assisted verification, autonomous visual reasoning, longer reasoning durations, and fine-tuning on an easier dataset, provide negligible improvement, indicating that this failure mode cannot be addressed through current prompting or post-hoc correction techniques. Our findings highlight a fundamental limitation of current VLM for inspection and monitoring tasks and underscore the need for architectural or training-level interventions that enable models to override internal expectations when confronted with contradictory evidence.
Wenqi Marshall Guo, Qingyun Qian, Shiyu Zhou +2
Department of CMPS, University of British Columbia, Canada · Weathon Software, Canada · Department of Computer Science, University of British Columbia, Canada
Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number, narrowing the inter-model gap on saturated suites and pressing models into low-score bands on harder ones. We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks. Across compact VLMs from the past two years, PScore separates model pairs more reliably than prior single-axis benchmarks under an item-level paired bootstrap, and surfaces behavioral differences these benchmarks average away. Even models with statistically indistinguishable PScores diverge sharply along the per-axis profile, particularly on sycophancy, which is nearly orthogonal to single-prompt accuracy. We will release the full pipeline, prompts, and per-item annotations.
This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages. We employ several fine-tuned vision-language models as independent annotators and combine their span predictions through character-level majority voting, and additionally explore activation probes. The approach ranks first in three of four languages and places on the podium in every language and metric. Our analysis indicates that disagreement among diverse models tracks disagreement among human annotators.
Toqeer Ehsan, Nico Penttilä, Richard Schmidt +2
Reliable Intelligence Team, Physical AI, VTT Technical Research Centre of Finland