We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer from fly and mouse to human and zebrafish. Most models were at chance zero-shot; a few examples helped mainly the closed and largest open ones. LoRA on a few thousand labels brought open models level with specialist models. When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors. The project will be publicly available upon acceptance.
Figures & tables
Figure 1: Synapse detection example inputs of fly. Left: presence , the query point marked by the yellow circle (answer: yes). Right: polarity , markers A and B inside the two partner processes (answer: B).
zero-shot
4-shot
LoRA
Model
fly P
fly Pol
mouse P
mouse Pol
fly P
fly Pol
mouse P
mouse Pol
fly P
fly Pol
mouse P
mouse Pol
Qwen2.5-VL-3B
49.1
50.0
49.5
50.0
50.9
50.0
53.9
50.0
50.3
57.8
74.4
76.4
Qwen2.5-VL-7B
50.0
49.0
49.6
49.2
48.8
49.1
48.2
49.3
71.8
70.0
86.3
81.4
Qwen2.5-VL-32B
48.2
49.1
46.2
49.8
55.2
49.4
52.2
50.3
68.2
61.7
83.5
77.4
Qwen3-VL-2B
51.2
49.1
48.3
50.9
50.0
48.7
54.1
50.0
69.7
56.9
81.0
79.9
Qwen3-VL-4B
53.3
48.2
47.5
51.9
59.1
51.7
48.8
49.8
72.4
56.3
83.7
82.1
Table 1: Synapse detection, balanced accuracy (%) on the test splits. P: presence ; Pol: polarity . Bold: best VLM per column; underlined : second best; Italic rows : specialist models, not ranked; –: not applicable.
Figure 2: Proofreading example inputs of fly. (a) A split-error site: the two segments are tinted red and blue; the proofreaders merged them, so the answer is yes. (b) A merge-error site: the segment is tinted red; the proofreaders split it, so the answer is no.
zero-shot
4-shot
LoRA
Model
S
M
S
M
S
M
Qwen3.5-9B
50.0
52.9
58.7
57.3
94.6
90.2
InternVL3-8B
51.0
50.0
49.0
57.3
93.9
90.1
Gemma 4 12B
50.1
50.0
62.7
58.0
95.0
90.9
CB2 ViT-B, EM only
–
–
–
–
94.2
91.1
Table 2: Proofreading on the complete ConnectomeBench2, balanced accuracy (%) on the test split: all 101,929 samples across 4 species. S: split-error subtask, M: merge-error subtask. Bold: best VLM per column; Italic row : the dataset’s own reported EM-only ViT-B; –: not applicable.
Qwen3.5-9B
InternVL3-8B
Gemma 4 12B
GPT-5.6 Sol
Claude Opus 5
ResNet-50
ConvNeXt-T
ViT-S/14
zero-shot
mouse S
50.0
53.4
50.4
67.2
73.1
–
–
–
mouse M
59.3
50.0
50.0
67.6
67.3
–
–
–
fly S
50.0
52.0
50.0
63.0
64.7
–
–
–
fly M
51.0
50.0
50.0
57.0
57.0
–
–
–
4-shot
Table 3: Proofreading on the subset of ConnectomeBench2 (Sec. 2.2 ), balanced accuracy (%). S: split-error subtask, M: merge-error subtask. Bold: best VLM per row; italic columns : specialist models, not ranked; –: not applicable.
Department of CMPS, University of British Columbia, Canada · Weathon Software, Canada · Department of Computer Science, University of British Columbia, Canada