A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care
Authors: Junseob Kim, Jade Chng, Ayman Ali, Victor Moas, Yichun Lee, Po-Chun Chin, Sunil Hwang, Rishikesan Kamaleswaran
Organizations: Department of Electrical and Computer Engineering, Duke University, Durham, NC, USA · Department of Biomedical Engineering, Duke University, Durham, NC, USA · Department of Surgery, Duke University, Durham, NC, USA · School of Medicine, College of Medicine, National Yang Ming Chiao Tung University, Taipei, Taiwan · Department of Mathematics, Korea Military Academy, Seoul, South Korea · Department of Anesthesiology, Duke University, Durham, NC, USA
Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constructed from public instructional and field TC3 videos and authoritative TC3 documents. It contains 581 items spanning 11 concepts, with 1,860 questions covering intervention recognition, doctrine, clinical reasoning, procedural guidance, and refusal when visual information is insufficient. Doctrine-based answers preserve verbatim source passages and character offsets. Construction combines visual annotation, passage retrieval, entailment checks, and verification across model families. Equipment boxes, anatomical labels, temporal segments, and source metadata accompany the question-answer pairs. Automated audits and ratings by two physicians and two medical students characterize annotation quality, with human ratings available for 88 retained items. The dataset provides a resource for adapting vision-language models to TC3, studying the connection between visual evidence and clinical knowledge, and evaluating recognition, doctrine recall, and abstention.
Figures & tables
Figure 1: Example TC3-VQA item ( ans_01332 ) from public-domain footage of the Committee on Tactical Combat Casualty Care (CoTCCC). Left: one of four frames with equipment, answer-region, and anatomy annotations. Right: four answerable question types with cited doctrine spans. The paired refusal question asks for the number of windlass turns, which the available frames do not show.
Figure 2: TC3-VQA construction workflow. Left: direct generation of questions and answers from an image. Right: visual concept assignment, doctrine retrieval, verbatim answer anchoring, and entailment filtering. Equipment and anatomy annotations record visual evidence; cross-family review and recognition consensus check concept labels. Refusal questions target information absent from the frames.
Figure 3: Example frames across the MARCH categories: (a) massive hemorrhage, (b) airway, (c) respiration, (d) circulation, and (e) hypothermia; IV/IO, intravenous or intraosseous. Frames are from public-domain CoTCCC footage. Overlays show released annotations from the 17 -class equipment ontology.
Figure 4: Composition of the released layers. (a) Information requested by the 150 refusal questions. (b) Source document of the cited answer, by question type. (c) Equipment-box counts for the nine most frequent of the 16 classes present. (d) Observed body region of the 431 answerable items.
MARCH
Concept
Answerable
Refusal
Safety-critical
M (massive hemorrhage)
Wound packing
129
28
129
Limb tourniquet application
137
55
137
Junctional hemorrhage control
4
2
4
A (airway)
Nasopharyngeal airway (NPA) insertion
41
17
2
Surgical cricothyroidotomy
11
5
11
R (respiration)
Chest seal application
29
11
29
Table 1: Released items by concept and MARCH category. Safety-critical counts answerable items whose cited doctrine passage the stakes rubric flags as safety-critical; the risk-weighted hallucination rate (RWHR; Usage Notes) uses concept-level weights. The construction vocabulary also includes tourniquet conversion, which has no released item. NPA, nasopharyngeal airway; TXA, tranexamic acid.
Judgment ( n )
Qwen3-VL
InternVL3
MedGemma
Majority
AC1
Refusal unanswerable ( 150 )
98.7 ( 12.2 )
99.3 ( 34.5 )
88.0 ( 3.6 )
98.7 ( 10.8 )
0.91
Doctrine question anchored ( 427 )
53.9 ( 4.7 )
50.4 ( 6.6 )
63.5 ( 52.7 )
58.8 ( 7.7 )
0.46
how question anchored ( 427 )
87.8 ( 4.9 )
93.4 ( 9.1 )
94.1 ( 89.2 )
94.6 ( 11.0 )
0.86
Reasoning answer correct ( 425 )
86.4 ( 30.8 )
89.4 ( 33.2 )
90.8 ( 38.0 )
91.5 ( 33.0 )
0.84
Table 2: Automated audit of released questions and answers. Values are percentages satisfying each criterion; parentheses give the corresponding control rate. A majority requires two of three votes, and Gwet’s AC1 measures inter-rater agreement. Anchored indicates that a question was judged specific to its paired frames.
Physicians (%)
Medical students (%)
Agreement (AC1)
Question type
Rater 1
Rater 2 †
Rater 1
Rater 2
Physicians
Students
All four
Recognition
69/20/10
69/23/8
93/3/3
94/1/5
0.80
0.95
0.75
Doctrine
74/17/9
91/3/6
95/3/1
76/17/7
0.66
0.71
0.68
Reasoning
59/30/11
92/1/7
92/6/2
60/33/7
0.48
0.52
0.57
Procedural ( how )
68/28/3
93/6/1
93/7/0
41/50/9
0.61
0.26
0.47
Table 3: Human ratings on the 88 adjudicated items retained in the release. Cells report correct / unclear / wrong for recognition and correct / minor problem / incorrect for quoted answers. AC1 is Gwet’s agreement within each rater pair and among all four raters.
Model
Recognition
Doctrine-MCQ
Doctrine-open
Refusal
RWHR ↓
easy
hard
entailed
unsupported ↓
Qwen2-VL-7B
0.854
0.951
0.864
0.267
0.576
0.260
0.921
Qwen2.5-VL-7B
0.805
0.963
0.874
0.285
0.574
0.833
0.546
Phi-3.5-Vision
0.666
0.867
0.703
0.225
0.588
0.933
0.888
InternVL3-8B
0.840
0.948
0.845
0.243
0.579
0.860
0.479
InternVL3-38B
0.740
0.986
0.913
0.264
0.523
0.933
0.420
Table 4: Baseline performance on 431 answerable and 150 refusal items. Doctrine-MCQ (multiple-choice questions) uses distractors from other concepts (easy) or other facets of the same concept (hard). Doctrine-open reports reference-sentence entailment and the proportion of response sentences unsupported by the reference or cited passage. Bold indicates the best value in each column; lower is better for unsupported and RWHR.
Figure 5: Recognition, abstention, and error across five baseline models. (a) Recognition and refusal accuracy with Wilson 95% intervals; marker area represents RWHR. (b) Coverage versus unweighted error (open markers) and severity-weighted error (filled markers).
Frame substituted (all 431 )
First frame only ( 204 multi-frame)
Model
own
swapped
Δ
abstain
window
1 frame
Δ
p
Qwen2-VL-7B
0.854
0.139
−0.715
2.3%
0.853
0.828
−0.025
0.42
Qwen2.5-VL-7B
0.807
0.081
−0.726
18.8%
0.814
0.686
−0.127
<0.001
Phi-3.5-Vision
0.666
0.332
−0.334
10.2%
0.716
0.637
−0.078
0.014
InternVL3-8B
0.840
0.123
−0.717
17.9%
0.873
0.745
−0.127
<0.001
InternVL3-38B
0.740
0.044
−0.696
44.8%
0.819
0.725
−0.093
0.002
Table 5: Recognition under frame substitution and temporal truncation. Left, questions paired with frames from another concept ( 431 items; uniform-choice reference 0.20 ). Right, full windows compared with the first frame alone ( 204 multi-frame items). Δ denotes perturbed minus original accuracy; p values are from two-sided exact McNemar tests.
Figure 6: Reference question-answer (QA) pairs compared with direct generation from the same frames. (a–c) Public-domain examples showing the pipeline-generated reference pair (Ours) and direct generation by Qwen2.5-VL-72B (Conventional). (d) Concept inconsistency during direct question generation (left) and closed-set concept selection (right) on 288 audit-confirmed candidates (Pixtral-Large [ 45 ] and the models of Sec. 2.4 ), with Wilson 95% intervals.