Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires current diagnostic criteria, standardized acquisition and reporting knowledge, longitudinal monitoring concepts, lesion morphology, and recognition of difficult mimics. We present MS-Exam-Gen, a reproducible framework for constructing and auditing a text-based multiple-choice question (MCQ) benchmark for MS-MRI knowledge; it does not evaluate direct MRI image interpretation. MS-Exam-Gen targets source-grounded criteria, protocols, reporting, and differential diagnosis. The framework combines expert-source indexing, exam-oriented topic induction, evidence-grounded MCQ generation, automated quality audits, a same-family consistency screen, and empirical calibration. From a 66-source corpus indexed into 4,289 retrieval chunks, the pipeline produced a locked 3,058-item candidate benchmark spanning 16 topics and 53 subtopics. Evaluation across 12 primary LLM endpoints yielded 36,696 item-level predictions and separated performance over a 42.8-percentage-point accuracy range (89.7% to 46.9%). Across these endpoints, 25.5% of items were missed by at least four. Post-generation audits showed that refreshed construction reduced measurable answer cues, while option-order testing showed that absolute MCQ scores remain position-sensitive. Generated construction labels remain metadata rather than validated psychometric categories. Because expert adjudication and full option-order counterbalancing remain future work, MS-Exam-Gen is not a clinically certified examination. It should be interpreted as an automatically filtered, source-grounded candidate benchmark and reproducible audit workflow for item-level and topic-specific LLM evaluation.
Scalable, but source grounding and clinical safeguards are variable.
✗
Variable
Variable
MS-Exam-Gen
This work
Text MCQ construction
Candidate benchmark only; expert adjudication and full option-order counterbalancing remain future work.
✓
✓
✓
TABLE I: Positioning of MS-Exam-Gen relative to related benchmark families.
Fig. 1: MS-Exam-Gen construction and audit workflow. Sources are indexed with metadata, sampled through seeded retrieval, organized into an exam-oriented taxonomy, used for source-grounded MCQ generation, and filtered by automated audits and a same-family consistency screen with regeneration of malformed or cueable items. Algorithm 1 additionally details benchmark locking, model evaluation, and empirical calibration. Construction labels are metadata; the locked bank is a candidate benchmark, not a clinical examination.
Stage
Input
Accepted/flagged
Locked
Standard Easy/Medium
650 each
646 each
646 each
Standard Hard
650
643
0
Refreshed Hard/synthesis
900 each
883 each
883 each
Legacy Hard/synthesis pool
1,293
1,217 flagged
0
TABLE II: Construction and audit trail for the locked BHI candidate benchmark.
Component
Value
Corpus/taxonomy
66 sources; 4,289 chunks; 16 topics; 53 subtopics
Locked bank
3,058 MCQs: Easy/Medium 646 each; Hard/synthesis 883 each
12 primary endpoints plus GPT-5 mini sensitivity; 36,696 primary/39,754 total responses; 0 parse failures
TABLE III: Final benchmark and evaluation summary.
Fig. 2: Benchmark auditing and empirical calibration. (A) The selector was fitted to the replaced legacy and locked Hard/synthesis strata. The dashed chance line applies only to the first two accuracy metrics; the third is item prevalence. Lower key predictability supports reduced measurable cueing. (B) Primary empirical difficulty is the number of 12 non-construction-overlap endpoints incorrect per item.
Fig. 3: Topic-level accuracy across 12 primary endpoints. All 16 submitted topic labels and item counts are retained. Blue brackets mark three semantically overlapping label pairs and report their accuracy-independent parent means. Cells report percentages, and the final column reports each original topic’s mean.
Model
Easy
Med.
Hard
Synth.
Avg.
95% CI
Gemini 2.5 Flash
91.0
88.4
88.1
91.3
89.7
86.0–92.9
Llama 4 Maverick
88.4
87.2
87.8
90.4
88.5
85.0–91.7
DeepSeek Chat
87.8
86.8
86.2
87.0
86.9
83.1–90.2
Claude Haiku 4.5
86.7
85.1
81.1
89.5
85.6
81.3–89.4
Gemini 2.5 Flash Lite
85.1
81.9
84.9
85.2
84.4
79.2–88.9
Qwen3 14B
83.3
82.5
83.5
83.8
83.3
78.3–87.8
TABLE IV: Primary construction-label-resolved model-panel results on the locked 3,058-item benchmark.
Department of Radiology, University of Calgary · Child and Adolescent Imaging Research (CAIR) Program · Alberta Children’s Hospital Research Institute +1
National Library of Medicine, Division of Intramural Research, Bethesda, MD, US · University of Illinois at Urbana-Champaign, Department of Computer Science, Urbana, IL, US · University of Michigan Medical School, Ann Arbor, Michigan, US