A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care
Authors: Junseob Kim, Jade Chng, Ayman Ali, Victor Moas, Yichun Lee, Po-Chun Chin, Sunil Hwang, Rishikesan Kamaleswaran
Organizations: Department of Electrical and Computer Engineering, Duke University, Durham, NC, USA · Department of Biomedical Engineering, Duke University, Durham, NC, USA · Department of Surgery, Duke University, Durham, NC, USA · School of Medicine, College of Medicine, National Yang Ming Chiao Tung University, Taipei, Taiwan · Department of Mathematics, Korea Military Academy, Seoul, South Korea · Department of Anesthesiology, Duke University, Durham, NC, USA
Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constructed from public instructional and field TC3 videos and authoritative TC3 documents. It contains 581 items spanning 11 concepts, with 1,860 questions covering intervention recognition, doctrine, clinical reasoning, procedural guidance, and refusal when visual information is insufficient. Doctrine-based answers preserve verbatim source passages and character offsets. Construction combines visual annotation, passage retrieval, entailment checks, and verification across model families. Equipment boxes, anatomical labels, temporal segments, and source metadata accompany the question-answer pairs. Automated audits and ratings by two physicians and two medical students characterize annotation quality, with human ratings available for 88 retained items. The dataset provides a resource for adapting vision-language models to TC3, studying the connection between visual evidence and clinical knowledge, and evaluating recognition, doctrine recall, and abstention.
Figures & tables
Figure 1: Example TC3-VQA item ( ans_01332 ) from public-domain footage of the Committee on Tactical Combat Casualty Care (CoTCCC). Left: one of four frames with equipment, answer-region, and anatomy annotations. Right: four answerable question types with cited doctrine spans. The paired refusal question asks for the number of windlass turns, which the available frames do not show.
Figure 2: TC3-VQA construction workflow. Left: direct generation of questions and answers from an image. Right: visual concept assignment, doctrine retrieval, verbatim answer anchoring, and entailment filtering. Equipment and anatomy annotations record visual evidence; cross-family review and recognition consensus check concept labels. Refusal questions target information absent from the frames.
Figure 3: Example frames across the MARCH categories: (a) massive hemorrhage, (b) airway, (c) respiration, (d) circulation, and (e) hypothermia; IV/IO, intravenous or intraosseous. Frames are from public-domain CoTCCC footage. Overlays show released annotations from the 17 -class equipment ontology.
Figure 4: Composition of the released layers. (a) Information requested by the 150 refusal questions. (b) Source document of the cited answer, by question type. (c) Equipment-box counts for the nine most frequent of the 16 classes present. (d) Observed body region of the 431 answerable items.
MARCH
Concept
Answerable
Refusal
Safety-critical
M (massive hemorrhage)
Wound packing
129
28
129
Limb tourniquet application
137
55
137
Junctional hemorrhage control
4
2
4
A (airway)
Nasopharyngeal airway (NPA) insertion
41
17
2
Surgical cricothyroidotomy
11
5
11
R (respiration)
Chest seal application
29
11
29
Table 1: Released items by concept and MARCH category. Safety-critical counts answerable items whose cited doctrine passage the stakes rubric flags as safety-critical; the risk-weighted hallucination rate (RWHR; Usage Notes) uses concept-level weights. The construction vocabulary also includes tourniquet conversion, which has no released item. NPA, nasopharyngeal airway; TXA, tranexamic acid.
Judgment ( n )
Qwen3-VL
InternVL3
MedGemma
Majority
AC1
Refusal unanswerable ( 150 )
98.7 ( 12.2 )
99.3 ( 34.5 )
88.0 ( 3.6 )
98.7 ( 10.8 )
0.91
Doctrine question anchored ( 427 )
53.9 ( 4.7 )
50.4 ( 6.6 )
63.5 ( 52.7 )
58.8 ( 7.7 )
0.46
how question anchored ( 427 )
87.8 ( 4.9 )
93.4 ( 9.1 )
94.1 ( 89.2 )
94.6 ( 11.0 )
0.86
Reasoning answer correct ( 425 )
86.4 ( 30.8 )
89.4 ( 33.2 )
90.8 ( 38.0 )
91.5 ( 33.0 )
0.84
Table 2: Automated audit of released questions and answers. Values are percentages satisfying each criterion; parentheses give the corresponding control rate. A majority requires two of three votes, and Gwet’s AC1 measures inter-rater agreement. Anchored indicates that a question was judged specific to its paired frames.
Physicians (%)
Medical students (%)
Agreement (AC1)
Question type
Rater 1
Rater 2 †
Rater 1
Rater 2
Physicians
Students
All four
Recognition
69/20/10
69/23/8
93/3/3
94/1/5
0.80
0.95
0.75
Doctrine
74/17/9
91/3/6
95/3/1
76/17/7
0.66
0.71
0.68
Reasoning
59/30/11
92/1/7
92/6/2
60/33/7
0.48
0.52
0.57
Procedural ( how )
68/28/3
93/6/1
93/7/0
41/50/9
0.61
0.26
0.47
Table 3: Human ratings on the 88 adjudicated items retained in the release. Cells report correct / unclear / wrong for recognition and correct / minor problem / incorrect for quoted answers. AC1 is Gwet’s agreement within each rater pair and among all four raters.
Model
Recognition
Doctrine-MCQ
Doctrine-open
Refusal
RWHR ↓
easy
hard
entailed
unsupported ↓
Qwen2-VL-7B
0.854
0.951
0.864
0.267
0.576
0.260
0.921
Qwen2.5-VL-7B
0.805
0.963
0.874
0.285
0.574
0.833
0.546
Phi-3.5-Vision
0.666
0.867
0.703
0.225
0.588
0.933
0.888
InternVL3-8B
0.840
0.948
0.845
0.243
0.579
0.860
0.479
InternVL3-38B
0.740
0.986
0.913
0.264
0.523
0.933
0.420
Table 4: Baseline performance on 431 answerable and 150 refusal items. Doctrine-MCQ (multiple-choice questions) uses distractors from other concepts (easy) or other facets of the same concept (hard). Doctrine-open reports reference-sentence entailment and the proportion of response sentences unsupported by the reference or cited passage. Bold indicates the best value in each column; lower is better for unsupported and RWHR.
Figure 5: Recognition, abstention, and error across five baseline models. (a) Recognition and refusal accuracy with Wilson 95% intervals; marker area represents RWHR. (b) Coverage versus unweighted error (open markers) and severity-weighted error (filled markers).
Frame substituted (all 431 )
First frame only ( 204 multi-frame)
Model
own
swapped
Δ
abstain
window
1 frame
Δ
p
Qwen2-VL-7B
0.854
0.139
−0.715
2.3%
0.853
0.828
−0.025
0.42
Qwen2.5-VL-7B
0.807
0.081
−0.726
18.8%
0.814
0.686
−0.127
<0.001
Phi-3.5-Vision
0.666
0.332
−0.334
10.2%
0.716
0.637
−0.078
0.014
InternVL3-8B
0.840
0.123
−0.717
17.9%
0.873
0.745
−0.127
<0.001
InternVL3-38B
0.740
0.044
−0.696
44.8%
0.819
0.725
−0.093
0.002
Table 5: Recognition under frame substitution and temporal truncation. Left, questions paired with frames from another concept ( 431 items; uniform-choice reference 0.20 ). Right, full windows compared with the first frame alone ( 204 multi-frame items). Δ denotes perturbed minus original accuracy; p values are from two-sided exact McNemar tests.
Figure 6: Reference question-answer (QA) pairs compared with direct generation from the same frames. (a–c) Public-domain examples showing the pipeline-generated reference pair (Ours) and direct generation by Qwen2.5-VL-72B (Conventional). (d) Concept inconsistency during direct question generation (left) and closed-set concept selection (right) on 288 audit-confirmed candidates (Pixtral-Large [ 45 ] and the models of Sec. 2.4 ), with Wilson 95% intervals.
Purpose: Vision-language models (VLMs) have shown promising performance in surgical visual question answering (VQA). However, existing surgical VQA datasets often contain linguistic shortcuts, where question phrasing implicitly constrains the answer space. It remains unclear whether reported performance reflects visual understanding or reliance on such linguistic shortcuts. Methods: We introduce SurgCheck, a diagnostic benchmark for quantifying linguistic shortcut reliance in surgical VQA. SurgCheck employs a paired-question design in which each surgical frame is associated with an original question containing entity names and a less-biased counterpart that removes these names while preserving identical visual content and ground-truth answers. The resulting performance gap provides a diagnostic signal of shortcut reliance. To ensure that the less-biased question remains well-defined even without entity names, four grounding cues are incorporated: bounding box, arrow, spatial position, and periphrasis. We evaluate both general-purpose and surgical-specific VLMs under zero-shot and fine-tuned settings on SurgCheck. To evaluate open-ended zero-shot responses, we introduce an LLM-as-a-judge evaluation protocol. Results: Using SurgCheck, we observe consistent performance degradation on less-biased questions across five VLMs, despite identical visual inputs. Text-only ablation reveals minimal performance drops for action and target prediction, indicating that action and target prediction is largely driven by linguistic shortcuts rather than visual reasoning. Conclusion: SurgCheck provides a controlled diagnostic framework that exposes failure modes masked by linguistic bias in existing surgical VQA benchmarks. Our findings demonstrate that strong benchmark performance does not necessarily imply faithful visual understanding, underscoring the need for bias-aware evaluation in surgical VQA.
Jongmin Shin, Ka Young Kim, Eunki Cho +2
Department of Surgery, Samsung Medical Center, Seoul 06351, Republic of Korea · Kyung Hee University, Yongin 17104, Republic of Korea
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves 79.0% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.
Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik +3
Adam Mickiewicz University · ARAAI Poland · NASK National Research Institute +2
Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, existing medical visual question answering (VQA) benchmarks collapse model capabilities into a single accuracy score, obscuring where and why models fail. We propose DeepTumorVQA, a hierarchical benchmark that follows the multi-stage evidence chain in tumor diagnosis and decomposes 3D CT reasoning into four stages: recognition, measurement, visual reasoning, and medical reasoning. Higher-level questions remain independently scorable, while their ground-truth evidence chains are defined over lower-level primitives. The benchmark contains 476K questions across 42 clinical subtypes on 9,262 3D CT volumes. In addition to a direct reasoning mode for VLMs, DeepTumorVQA provides tool-interaction environments for agent evaluation, where a model can call external tools, including segmentation models, measurement programs, and medical knowledge modules, before answering the question. Evaluating over 30 model configurations, we find that reliable quantitative measurement is the primary bottleneck, making later-stage visual and medical reasoning harder for VLMs, while tool augmentation substantially mitigates this issue. When tools are available, leveraging medical knowledge and tools to reason about medical images becomes a new challenge. We further show that ground-truth step-by-step tool-use traces from DeepTumorVQA can supervise agents and reduce tool-use and reasoning failures. This stage-wise progression from recognition to measurement to visual and medical reasoning provides a concrete roadmap for future medical VLM and AI agent studies. All data and code are released at https://github.com/Schuture/DeepTumorVQA.
Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi +7
Johns Hopkins University · University of Bologna · Center for Biomolecular Nanotechnologies, Istituto Italiano di Tecnologia +3