templar: agentic induction and evolution of standardized radiology reporting templates from large-scale clinical corpora
Organizations: Tsinghua University · University of Hong Kong · University of California, San Diego
Abstract
Structured radiology reporting mitigates the heterogeneity of free-text reports, yet its benefits depend on high-quality reporting templates. In practice, such templates are conventionally built through labor-intensive expert consensus and therefore vary across institutions and lag behind evolving clinical practice. Large language models (LLMs) enable automated template induction, but existing approaches remain limited: single-LLM induction is constrained by context length, and the corpus-scale method ASTAR produces a static, closed-corpus template without external grounding or downstream adaptation. To address these limitations, we propose TEMPLAR, a TEMPLate-centric Agentic framework for inducing and evolving standardized Radiology reporting templates from large-scale clinical corpora. TEMPLAR treats the template as a persistent central state maintained alongside two provenance-aware knowledge graphs, namely an anatomical graph that constrains template construction and a diagnostic graph that supports finding-to-diagnosis reasoning. Three agents operate on this state. The Induction Agent derives canonical clinical slots from anatomy-constrained Span-Triple atoms via dual-view similarity clustering; the Evolution Agent then assembles these slots into a hierarchical template and revises it under consistency constraints, external clinical evidence, and downstream structuring feedback; and the Clinical Agent applies the evolved template to report structuring, reconstruction, and diagnostic reasoning. Across four datasets, TEMPLAR outperforms ASTAR, three medical LLMs, and six general-purpose LLMs in coverage, information fidelity, and diagnostic fidelity, while achieving the highest or tied-highest LLM-rated template quality. Its fidelity advantages over ASTAR persist under cross-dataset transfer, and cumulative ablations support complementary contributions of its key components.
Figures & tables
| Template Coverage | Information Fidelity | Diagnostic Fidelity | ||||||||||||
| Method | Case | Key | Avg. | R-1 | R-2 | R-L | chrF | chrF++ | BERT | Avg. | PDA | KFP | CA | Avg. |
| PETCT ( ) | ||||||||||||||
| Medical LLM baselines | ||||||||||||||
| lingshu-32b | 0.6678 | 0.7192 | 0.6935 | 0.5920 | 0.3089 | 0.3937 | 0.4916 | 0.4687 | 0.8865 | 0.5236 | 0.4939 | 0.5310 | 0.6000 | 0.5416 |
| medgemma-27b | 0.6744 | 0.7027 | 0.6886 | 0.7495 | 0.5163 | 0.5727 | 0.6214 | 0.6058 | 0.9069 | 0.6621 | 0.5633 | 0.6051 | 0.6571 | 0.6085 |
| clinfusion-32b | 0.4710 | 0.4910 | 0.4810 | 0.1419 | 0.0580 | 0.1089 | 0.0598 | 0.0556 | 0.7380 | 0.1937 | 0.5102 | 0.5320 | 0.6204 | 0.5542 |
| Template Coverage | Information Fidelity | Diagnostic Fidelity | ||||||||||||
| Method | Case | Key | Avg. | R-1 | R-2 | R-L | chrF | chrF++ | BERT | Avg. | PDA | KFP | CA | Avg. |
| MERLIN ( ) | ||||||||||||||
| Medical LLM baselines | ||||||||||||||
| lingshu-32b | 0.5214 | 0.5347 | 0.5281 | 0.5677 | 0.3047 | 0.3765 | 0.4965 | 0.4701 | 0.8521 | 0.5113 | 0.8662 | 0.8975 | 0.8637 | 0.8758 |
| medgemma-27b | 0.3230 | 0.3801 | 0.3516 | 0.6072 | 0.3891 | 0.3860 | 0.4705 | 0.4538 | 0.8766 | 0.5305 | 0.7137 | 0.8057 | 0.7456 | 0.7550 |
| clinfusion-32b | 0.3293 | 0.3435 | 0.3364 | 0.2769 | 0.1159 | 0.1589 | 0.2084 | 0.1947 | 0.7833 | 0.2897 | 0.5853 | 0.6650 | 0.6333 | 0.6279 |
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | ID | Evaluation criterion |
|---|---|---|
| Coverage and completeness | TD1 | Coverage of the essential structures and measurements required in PET/CT reports. |
| Coverage and completeness | TD2 | Explicit support for common abnormalities, including primary tumors, lymph-node metastases, and distant metastases involving the bone, liver, or lung. |
| Structure and granularity | TD3 | Hierarchical organization consistent with radiological reasoning, proceeding from findings to organ systems and key anatomical structures. |
| Structure and granularity | TD4 | Appropriate field granularity that preserves clinically relevant information without making the template unnecessarily fragmented or burdensome to complete. |
| Terminology and clarity | TD5 | Clear field names and Chinese labels that conform to commonly used clinical terminology. |
| Terminology and clarity | TD6 | Natural and usable permissible values, such as Normal , Abnormal , Uncertain , and NotAssessed . |
| Category | ID | Evaluation criterion |
|---|---|---|
| Omission reduction | CU1 | The template reduces omissions of clinically important findings. |
| Reporting consistency | CU2 | The template improves consistency across PET/CT reports. |
| Workflow alignment | CU3 | The ordering of template fields is consistent with the clinician’s image-reading workflow. |
| Unassessed-state support | CU4 | The template enables unassessed structures or findings to be represented naturally using the NotAssessed state. |
| Communication efficiency | CU5 | The template improves communication efficiency between radiologists and referring clinicians. |
| Overall satisfaction | CU6 | Overall evaluator satisfaction with the template. |
| Template Coverage | Information Fidelity | Diagnostic Fidelity | ||||||||||||
| Method | Case | Key | Avg. | R-1 | R-2 | R-L | chrF | chrF++ | BERT | Avg. | PDA | KFP | CA | Avg. |
| Medical LLM baselines | ||||||||||||||
| lingshu-32b | 0.5213 | 0.5339 | 0.5276 | 0.5171 | 0.2750 | 0.3669 | 0.3902 | 0.3742 | 0.8770 | 0.4667 | 0.4861 | 0.5172 | 0.5304 | 0.5112 |
| medgemma-27b | 0.5388 | 0.5548 | 0.5468 | 0.6382 | 0.4104 | 0.4717 | 0.4873 | 0.4762 | 0.8970 | 0.5634 | 0.5710 | 0.6328 | 0.6232 | 0.6090 |
| clinfusion-32b | 0.4894 | 0.4930 | 0.4912 | 0.3417 | 0.1849 | 0.2407 | 0.2180 | 0.2125 | 0.7995 | 0.3329 | 0.6236 | 0.6699 | 0.6873 | 0.6603 |
| Medical LLM Avg. | 0.5165 | 0.5272 | 0.5219 | 0.4990 | 0.2901 | 0.3598 | 0.3651 | 0.3543 | 0.8578 | 0.4543 | 0.5602 | 0.6066 | 0.6136 | 0.5935 |
| Model | Finish reason | Output tokens | Outcome |
|---|---|---|---|
| seed-2.0-pro | length | 21,623 | truncated JSON |
| seed-1.8 | length | 18,836 | truncated JSON |
| seed-2.0-mini | length | 18,634 | truncated JSON |
| seed-2.0-lite | length | 11,541 | truncated JSON |
| seed-1.6 | stop | 6,982 | valid JSON in screening |
| seed-2.1-pro | – | – | no response after approximately 40 min |
| Dataset | Peak input | Peak output | Sum of peaks |
|---|---|---|---|
| PET/CT | 24,872 | 13,332 | 38,204 |
| ViMed | 29,387 | 8,964 | 38,351 |
| BDMAP | 39,461 | 3,558 | 43,019 |
| MERLIN | 30,039 | 13,585 | 43,624 |
| Cause | Count | Structural signature |
|---|---|---|
| Truncation | 78 | 1–59 missing } and up to 14 missing ] |
| Generation restart | 17 | a second JSON fence or a second top-level desc |
| Excess closing delimiter | 17 | an otherwise complete object followed by extra delimiters |
| Balanced but invalid syntax | 3 | e.g., an unambiguous trailing comma |
| Dataset | Examples | Categorical | Numerical |
|---|---|---|---|
| PETCT | 12 | 11 | 1 |
| ViMed | 12 | 11 | 1 |
| MERLIN | 14 | 13 | 1 |
| BDMAP | 12 | 8 | 4 |
| Total | 50 | 43 | 7 |