MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge
Organizations: eBRAIN Lab, Division of Engineering, New York University Abu Dhabi (NYUAD), Abu Dhabi, UAE
Abstract
Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires current diagnostic criteria, standardized acquisition and reporting knowledge, longitudinal monitoring concepts, lesion morphology, and recognition of difficult mimics. We present MS-Exam-Gen, a reproducible framework for constructing and auditing a text-based multiple-choice question (MCQ) benchmark for MS-MRI knowledge; it does not evaluate direct MRI image interpretation. MS-Exam-Gen targets source-grounded criteria, protocols, reporting, and differential diagnosis. The framework combines expert-source indexing, exam-oriented topic induction, evidence-grounded MCQ generation, automated quality audits, a same-family consistency screen, and empirical calibration. From a 66-source corpus indexed into 4,289 retrieval chunks, the pipeline produced a locked 3,058-item candidate benchmark spanning 16 topics and 53 subtopics. Evaluation across 12 primary LLM endpoints yielded 36,696 item-level predictions and separated performance over a 42.8-percentage-point accuracy range (89.7% to 46.9%). Across these endpoints, 25.5% of items were missed by at least four. Post-generation audits showed that refreshed construction reduced measurable answer cues, while option-order testing showed that absolute MCQ scores remain position-sensitive. Generated construction labels remain metadata rather than validated psychometric categories. Because expert adjudication and full option-order counterbalancing remain future work, MS-Exam-Gen is not a clinically certified examination. It should be interpreted as an automatically filtered, source-grounded candidate benchmark and reproducible audit workflow for item-level and topic-specific LLM evaluation.
Figures & tables
| Family | Representative examples | Useful for | Limitation for this study | MS-MRI text | MS-MRI evidence | MCQ audit / calib. |
|---|---|---|---|---|---|---|
| Broad medical QA | MedQA, PubMedQA, MMLU, MedMCQA, and MultiMedQA [ 5 , 6 , 7 , 16 , 2 ] | Broad text/exam QA | Not organized around MS-MRI criteria, protocols, reporting, or mimics. | ✗ | Partial | ✗ |
| Clinical/multimodal suites | MultiMedQA-style and broader medical benchmark suites [ 2 , 17 ] | Clinical QA and scenarios | Coverage and provenance vary; usually not a locked MS-MRI source-grounded MCQ bank. | ✗ | Variable | Variable |
| Radiology VQA/report | Radiology foundation-model and report-generation benchmarks [ 18 , 19 ] | Image, VQA, reports | Image/report focused; not designed for guideline-source textual MCQ auditing. | ✗ | Partial | Variable |
| MS imaging tasks | MSSEG and longitudinal lesion segmentation [ 11 , 12 ] | Lesion segmentation | Domain-specific imaging tasks, but image labels/masks rather than textual source evidence. | ✗ | Partial | ✗ |
| LLM-generated evals | Model-written, MT-Bench, G-Eval, LLM-as-judge methods; MCQ-bias analyses [ 13 , 14 , 15 , 20 , 21 ] | Generation and judging | Scalable, but source grounding and clinical safeguards are variable. | ✗ | Variable | Variable |
| MS-Exam-Gen | This work | Text MCQ construction | Candidate benchmark only; expert adjudication and full option-order counterbalancing remain future work. | ✓ | ✓ | ✓ |
| Stage | Input | Accepted/flagged | Locked |
|---|---|---|---|
| Standard Easy/Medium | 650 each | 646 each | 646 each |
| Standard Hard | 650 | 643 | 0 |
| Refreshed Hard/synthesis | 900 each | 883 each | 883 each |
| Legacy Hard/synthesis pool | 1,293 | 1,217 flagged | 0 |
| Component | Value |
|---|---|
| Corpus/taxonomy | 66 sources; 4,289 chunks; 16 topics; 53 subtopics |
| Locked bank | 3,058 MCQs: Easy/Medium 646 each; Hard/synthesis 883 each |
| Final audit | 0 schema errors; 0 exact duplicates; 54 retained near-duplicate transparency flags |
| Evaluation | 12 primary endpoints plus GPT-5 mini sensitivity; 36,696 primary/39,754 total responses; 0 parse failures |
| Model | Easy | Med. | Hard | Synth. | Avg. | 95% CI |
|---|---|---|---|---|---|---|
| Gemini 2.5 Flash | 91.0 | 88.4 | 88.1 | 91.3 | 89.7 | 86.0–92.9 |
| Llama 4 Maverick | 88.4 | 87.2 | 87.8 | 90.4 | 88.5 | 85.0–91.7 |
| DeepSeek Chat | 87.8 | 86.8 | 86.2 | 87.0 | 86.9 | 83.1–90.2 |
| Claude Haiku 4.5 | 86.7 | 85.1 | 81.1 | 89.5 | 85.6 | 81.3–89.4 |
| Gemini 2.5 Flash Lite | 85.1 | 81.9 | 84.9 | 85.2 | 84.4 | 79.2–88.9 |
| Qwen3 14B | 83.3 | 82.5 | 83.5 | 83.8 | 83.3 | 78.3–87.8 |