MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction
Organizations: Beijing University of Posts and Telecommunications · West China Hospital · Sichuan Provincial Engineering Research Center of Intelligent Diagnosis and Treatment of Breast Diseases
Abstract
Large-scale construction of medical vision-language model (VLM) benchmarks is increasingly feasible with richly annotated imaging datasets and large language models (LLMs), yet existing automation largely focuses on generating evaluation items within predefined benchmark specifications. We study the broader problem of automatically deriving the specification itself: what to evaluate, which annotations support each task, and how to translate this evidence into reliable evaluation items. We formulate benchmark construction as constrained compilation, in which the benchmark specification is progressively derived from evaluation requirements, heterogeneous annotations, and medical knowledge. Based on this formulation, we introduce MedBenchAgent, a multi-agent framework with a Benchmark Intermediate Representation (BIR) that encodes task definitions, evidence mappings, evaluation protocols, and item specifications across construction stages. MedBenchAgent separates planning, which derives and verifies the specification, from instantiation, which constructs and audits items under the locked specification. MedBenchAgent achieves a Task-Space F1 of 90.9%, outperforming direct task induction (79.2-80.0%) and prior-guided induction (85.1%); 994 of 1,000 sampled items from correctly identified tasks pass human audit. We further demonstrate portability to a specialized medical domain and evaluate twelve VLMs, revealing task- and setting-specific variation obscured by aggregate scores. These results establish constrained compilation as a scalable and auditable framework for medical VLM benchmark construction beyond question generation.
Figures & tables
| Anatomy | Finding | Assessment & Diagnosis | |||||
| Model | AA | FT | FA | AC | DX | PS | Overall |
| General-purpose VLMs | |||||||
| InternVL3-8B | 36.59 / 17.30 | 19.09 / 8.65 | 42.28 / 30.86 | 16.12 / 12.82 | 35.35 / 22.86 | 48.95 / 40.25 | 33.06 / 22.12 |
| Qwen2.5-VL-7B | 48.86 / 23.08 | 29.77 / 10.17 | 41.45 / 29.28 | 21.35 / 11.93 | 33.76 / 22.67 | 51.03 / 41.86 | 37.70 / 23.16 |
| Qwen3-VL-8B-I | 34.89 / 18.56 | 16.31 / 9.03 | 29.84 / 22.57 | 39.11 / 19.74 | 39.98 / 29.65 | 60.55 / 47.85 | 36.78 / 24.57 |
| Qwen3-VL-8B-T † | 25.52 / 14.18 | 15.09 / 7.29 | 56.93 / 35.13 | 38.58 / 20.03 | 35.13 / 24.15 | 53.05 / 42.52 | 37.38 / 23.88 |
| VinDr-Mammo | BUS-BRA | ||
| Model | Det | VG | Det |
| General-purpose VLMs | |||
| InternVL3-8B | 3.30 / 0.24 | 3.15 / 0.33 | 26.00 / 11.20 |
| Qwen2.5-VL-7B | 7.60 / 3.91 | 8.77 / 6.51 | 28.94 / 22.24 |
| Qwen3-VL-8B-I | 10.72 / 7.00 | 13.68 / 10.98 | 51.11 / 56.40 |
| Qwen3-VL-8B-T † | 10.94 / 7.32 | 10.42 / 7.00 | 37.50 / 32.88 |
| Anatomy | Finding | Assessment & Diagnosis | Overall | Overall | ||||
| Model | AA | FT | FA | AC | DX | PS | Acc/F1 | Acc/F1 |
| Random Guess | 25.00 / 20.64 | 12.50 / 8.02 | 36.90 / 30.95 | 19.76 / 16.84 | 25.00 / 22.39 | 41.67 / 37.83 | 26.81 / 22.78 | – |
| MedGemma-4B | 11.56 / 11.58 | 8.24 / 5.88 | 39.06 / 27.99 | 1.83 / 3.38 | 21.06 / 16.73 | 40.44 / 36.59 | 20.36 / 17.02 | 12.48 / 6.97 |
| Lingshu-7B | 26.78 / 19.60 | 16.48 / 7.03 | 55.03 / 27.96 | 32.85 / 16.08 | 35.98 / 18.74 | 61.92 / 32.28 | 38.17 / 20.28 | 8.83 / 9.68 |
| Hulu-Med-7B | 19.06 / 16.62 | 5.71 / 4.69 | 54.83 / 28.07 | 31.07 / 15.89 | 38.27 / 22.01 | 52.23 / 41.12 | 33.53 / 21.40 | 5.86 / 6.33 |
| BreastGPT-8B | 36.66 / 18.00 | 42.01 / 11.48 | 58.54 / 28.81 | 20.12 / 11.37 | 30.70 / 18.49 | 35.55 / 21.97 | 37.26 / 18.35 | 17.97 / 22.61 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Raw proposed task | Raw supporting field(s) |
| BrEaST: 11 retained proposals | ||
| BrEaST | anatomy_attribute.tissue_composition | study.annotation.tissue_composition |
| BrEaST | finding_attribute.shape | image.annotation.shape |
| BrEaST | finding_attribute.margin | image.annotation.margin |
| BrEaST | finding_attribute.echogenicity | image.annotation.echogenicity |
| BrEaST | finding_attribute.posterior_features | image.annotation.posterior_features |
| Method | Dataset | Proposal or confirmed task | Label | Review rationale |
| MedBenchAgent | BUS-BRA | finding type | FP | Duplicated diagnosis on the histology answer axis. |
| MedBenchAgent | BUS-BRA | visual grounding | FP | Did not define a capability distinct from generic lesion detection. |
| MedBenchAgent | VinDr | view identification | FP | Burned-in view markers exposed the target as a shortcut. |
| MedBenchAgent | VinDr | finding-level assessment | FN | Finding-level BI-RADS supported a distinct, initially omitted answer axis. |
| Direct DeepSeek | BrEaST | clinical signs | FP | Nonvisual clinical metadata. |
| Direct DeepSeek | BrEaST | patient symptoms | FP | Nonvisual patient information. |
| Method | Dataset | Proposal or confirmed task | Label | Review rationale |
| Direct Qwen | BrEaST | palpability assessment | FP | Clinical palpation is not image-answerable. |
| Direct Qwen | BrEaST | interpretation classification | FP | A source conclusion is not an independent visual capability. |
| Direct Qwen | BrEaST | verification method | FP | Nonvisual confirmation metadata. |
| Direct Qwen | BrEaST | patient symptom status | FP | Nonvisual patient information. |
| Direct Qwen | BUS-BRA | lateralization | FP | Laterality metadata was outside the intended scope. |
| Direct Qwen | BUS-BRA | device identification | FP | Acquisition-device metadata was outside the intended scope. |
| Task | Supporting field | Resp. | Metric | Planned | Realized |
| Diagnosis | image.annotation.diagnosis | letter | Macro-F1 | 1,667 | 1,512 |
| Detection | image.annotation.findings[] .bbox[] | bbox | mIoU | 1,666 | 1,000 |
| Globules presence | image.annotation.attribute_presence .globules | letter | Accuracy | 334 | 334 |
| Milia-like cyst presence | image.annotation.attribute_presence .milia_like_cyst | letter | Macro-F1 | 334 | 334 |
| Negative network presence | image.annotation.attribute_presence .negative_network | letter | Macro-F1 | 333 | 333 |
| Pigment network presence | image.annotation.attribute_presence .pigment_network | letter | Accuracy | 333 | 333 |
| Finding Attribute | |||||||
| Model | Diagnosis | Globules | Milia-like | Neg. net. | Pigment net. | Streaks | Detection |
| General-purpose VLMs | |||||||
| InternVL3-8B | 20.7 / 15.3 | 31.4 / 31.1 | 70.7 / 56.6 | 45.3 / 35.0 | 72.4 / 63.4 | 39.3 / 35.7 | 50.5 / 52.2 |
| Qwen2.5-VL-7B | 24.2 / 14.6 | 57.8 / 53.7 | 80.8 / 55.5 | 83.8 / 47.4 | 59.5 / 59.3 | 33.6 / 30.7 | 75.1 / 86.6 |
| Qwen3-VL-8B-I | 53.2 / 22.8 | 28.4 / 27.5 | 80.2 / 61.8 | 80.2 / 47.3 | 59.5 / 58.4 | 58.0 / 44.7 | 80.3 / 91.4 |
| Qwen3-VL-8B-T † | 30.9 / 18.6 | 60.5 / 52.9 | 83.2 / 61.1 | 85.0 / 48.2 | 57.1 / 56.0 | 82.9 / 47.0 | 71.2 / 83.6 |
| Anatomy | Finding | Assessment & Diagnosis | |||
| Model | AA | FA | AC | DX | PS |
| General-purpose VLMs | |||||
| InternVL3-8B | 27.5 / 10.8 | 42.3 / 30.9 | 18.8 / 10.1 | 28.9 / 16.3 | 50.8 / 33.4 |
| Qwen2.5-VL-7B | 33.0 / 19.5 | 41.5 / 29.3 | 16.4 / 10.5 | 34.7 / 21.3 | 54.7 / 36.4 |
| Qwen3-VL-8B-I | 27.5 / 10.9 | 29.8 / 22.6 | 20.7 / 16.7 | 38.8 / 26.9 | 60.9 / 40.3 |
| Qwen3-VL-8B-T † | 28.0 / 13.4 | 56.9 / 35.1 | 21.1 / 19.3 | 38.0 / 24.5 | 49.2 / 32.7 |
| Finding Attribute | |||||||
| Model | Shape | Margin | Echo. | Posterior | Calc. | Halo | Skin |
| General-purpose VLMs | |||||||
| InternVL3-8B | 54.5 / 30.9 | 53.2 / 53.0 | 46.3 / 33.1 | 21.8 / 15.6 | 45.2 / 23.0 | 26.0 / 24.4 | 48.8 / 35.9 |
| Qwen2.5-VL-7B | 55.4 / 26.8 | 58.3 / 47.4 | 40.3 / 28.3 | 19.4 / 8.1 | 27.8 / 14.7 | 47.6 / 44.6 | 41.4 / 35.1 |
| Qwen3-VL-8B-I | 34.3 / 28.6 | 47.2 / 42.2 | 25.5 / 16.2 | 37.3 / 24.8 | 50.4 / 24.0 | 11.4 / 16.7 | 2.7 / 5.6 |
| Qwen3-VL-8B-T † | 56.6 / 34.4 | 58.3 / 45.8 | 61.5 / 28.4 | 22.2 / 12.4 | 59.9 / 30.4 | 63.4 / 46.6 | 76.6 / 47.8 |
| Assessment & Diagnosis | |||
| Model | AC | DX | PS |
| General-purpose VLMs | |||
| InternVL3-8B | 25.7 / 24.1 | 41.8 / 29.4 | 47.1 / 47.1 |
| Qwen2.5-VL-7B | 37.0 / 15.6 | 32.8 / 24.1 | 47.4 / 47.3 |
| Qwen3-VL-8B-I | 35.3 / 20.5 | 41.1 / 32.4 | 60.2 / 55.4 |
| Qwen3-VL-8B-T † | 31.2 / 20.9 | 32.2 / 23.8 | 56.9 / 52.3 |
| Anatomy | Finding | Assessment & Diagnosis | |
| Model | AA | FT | AC |
| General-purpose VLMs | |||
| InternVL3-8B | 45.7 / 23.8 | 19.1 / 8.7 | 3.9 / 4.2 |
| Qwen2.5-VL-7B | 64.7 / 26.6 | 29.8 / 10.2 | 10.7 / 9.6 |
| Qwen3-VL-8B-I | 42.3 / 26.2 | 16.3 / 9.0 | 61.4 / 22.0 |
| Qwen3-VL-8B-T † | 23.0 / 14.9 | 15.1 / 7.3 | 63.4 / 19.9 |
| BUS-BRA Detection | VinDr Detection | VinDr Visual Grounding | ||||
| Model | mIoU | [email protected] | mIoU | [email protected] | mIoU | [email protected] |
| General-purpose VLMs | ||||||
| InternVL3-8B | 26.0 | 11.2 | 3.3 | 0.2 | 3.2 | 0.3 |
| Qwen2.5-VL-7B | 28.9 | 22.2 | 7.6 | 3.9 | 8.8 | 6.5 |
| Qwen3-VL-8B-I | 51.1 | 56.4 | 10.7 | 7.0 | 13.7 | 11.0 |
| Qwen3-VL-8B-T † | 37.5 | 32.9 | 10.9 | 7.3 | 10.4 | 7.0 |