Symphony for Text Generation: Benchmarking Clinical Note Generation
Authors: Daniel Varab, Victor Petrén Bach Hansen, Asbjørn W. Helge, Kevin Pelgrims, Mathias Baltzersen, Adrian Young-San Roessler, Vanessa Klungtvedt, Maximilian Brand, +3 more
Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.
Figures & tables
Figure 1 : Pairwise preference outcomes on SOAP notes, shown for each dataset and overall pooled. For each Corti–comparator pair, bars report the percentages of judgments preferring Corti, preferring the comparison system, or resulting in a tie, aggregated over the eight adapted PDSQI dimensions. Contradictory judgments are counted as ties. The pooled result comprises 3,296 dimension-level judgments per compared system (412 encounters × 8 dimensions).
Dataset
N
#words (min–max)
Avg. encounter duration (min)
ACI-BENCH (EN)
112
1271 (710–2541)
8.5
MedConv (DA)
100
1655 (923–2585)
11.0
MedConv (DE)
100
1822 (904–2739)
12.1
MedConv (EN)
100
2031 (930–3073)
13.5
Table 1 : Dataset sizes, mean transcript length, and estimated durations at 150 words-per-minute.
Dataset
System
Groundedness ↑
Completeness ↑
Conciseness ↑
ACI-Bench (EN)
Corti
96.3% ± 0.3%
77.3% ± 0.2%
74.9% ± 0.2%
Heidi
94.9% ± 0.4%
75.9% ± 0.3%
75.7% ± 0.2%
Tandem Health
94.7% ± 0.3%
73.1% ± 0.4%
74.0% ± 0.1%
MedConv (EN)
Corti
97.8% ± 0.1%
73.0% ± 0.1%
81.4% ± 0.3%
Heidi
95.4% ± –
64.6% ± –
83.0% ± – 0.0%
Tandem Health
96.7% ± –
62.9% ± –
81.3% ± –
Table 2 : Content related metrics across datasets. Highest scores within each dataset are shown in bold. ± denotes standard deviation over means across multiple runs; – marks single-run entries.
Corti
Heidi
Tandem Health
Dataset
Transcript chars
avg (s)
avg (s)
vs Corti
avg (s)
vs Corti
ACI-BENCH (EN)
6101 (3474–12097)
5.86
15.59
2.7 ×
11.33
1.9 ×
MedConv (DA)
7591 (4243–12115)
9.68
19.68
2.0 ×
21.41
2.2 ×
MedConv (DE)
9551 (5269–15258)
9.07
19.22
2.1 ×
18.97
2.1 ×
MedConv (EN)
8960 (4973–13949)
7.49
17.60
2.3 ×
14.18
1.9 ×
Table 3 : Speedup ratios and dataset statistics relative to Corti.
Figure 2 : PDSQI-9 preference scores on SOAP, comparing Corti with each evaluated system.
Figure 3 : PDSQI-9 preference scores after adding general-practice instructions, comparing Corti with each evaluated system. Each dot represents a comparison between Corti and one other system.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Corti
Heidi
Tandem Health
Access method
Text Generation API
End-user web interface
End-user web interface
Input
Pre-transcribed consultation transcript; additional extracted clinical facts available to the generation pipeline
Pre-transcribed consultation transcript provided as context
Pre-transcribed consultation transcript provided as context
Speech-to-text evaluated
No
No
No
Markets / languages
UK English, German, Danish
UK English, German, Danish
UK English, German, Danish
Clinical note templates
SOAP and GP-SOAP configurations
SOAP configurations
SOAP configurations
Template source
Corti Standards and evaluation-specific configurations
Heidi defaults for each market, adapted where necessary
Tandem Health defaults where available; templates created from scratch where no suitable default was available
Appendix
Table 4 : Configuration of the three clinical note-generation systems evaluated in the benchmark. All systems were provided with pre-transcribed consultation content; speech-to-text performance was therefore outside the scope of the comparison.
Country
Default available
Adjustments
UK
Yes
Moved prompts of separate Past Medical History section into Subjective section to normalise to 4 SOAP sections
Germany
Yes
Adjusted German SOAP-Notiz by moving all of Past Medical History (Vergangene Krankengeschichte) under Subjective section to normalise to 4 SOAP sections. Removed (S), (O), (A), (P) from headings to normalise.
Denmark
Yes (English prompts)
Exact same adjustments as with UK SOAP
Appendix
Table 5 : Heidi SOAP template configuration by country.
Country
Default available
Adjustments
UK
No
Created from scratch, headings only
Germany
Yes
Based on standard German EHR template (SOP), separated “Beurteilung” (Assessment) from “Plan” to normalise to 4-section SOAP
Denmark
Yes
Based on SOAP for Almen Praksis, separated “Vurdering” (Assessment) from Plan to normalise to 4-section SOAP
Appendix
Table 6 : Tandem Health SOAP template configuration by country.
Property
Corti
Heidi
Tandem Health
Template representation
Typed, section-based templates composed of reusable document sections
Editable template structure with section headings and associated prompting
Primarily section- and heading-based templates
Content control
Explicit section-level instructions define content to include or exclude
Content instructions can be specified within the editable template prompt
Template headings define intended content; additional custom instructions can be added
Style and synthesis control
Separate controls for writing style, synthesis instructions, and output representation
Style and content instructions are expressed within the template prompt
Custom prompting and user- or clinic-level preferences can influence generated output
Output structure
Section-specific, schema-controlled outputs, including typed structured fields
Template structure determines the organization of the generated document
Template headings determine document structure; selected fields can use structured output types
Prompt transparency
Section-level content, style, synthesis, and formatting instructions are exposed
Template prompt can be inspected through the Structure view
System-provided section-level instructions are not fully exposed
Reuse and inheritance
Sections are reusable and versioned; templates can inherit from shared base sections
Vendor defaults can be copied, modified, or replaced with custom templates
Vendor defaults can be modified or new templates can be created
Appendix
Table 8 : Template architecture and configuration transparency across the evaluated systems.
Dimension
Description
Cited
Claims are accompanied by appropriate citations to the source documentation.
Accurate
Information is factually correct and free from fabrication, falsification, or other incorrect content.
Thorough
Important clinical information is sufficiently complete, without clinically relevant omissions.
Useful
Included information is relevant and useful to the clinician receiving the summary.
Organized
Information is structured coherently so that the clinical course and relevant information can be followed easily.
Comprehensible
The summary is clear and unambiguous and uses understandable language and terminology.
Appendix
Table 9 : Dimensions of the Provider Documentation Summarization Quality Instrument (PDSQI-9) Croxford et al. [2025b] .
Figure 4 : Judge sensitivity under GPT-5.4 and Opus 4.6 across all datasets and dimensions.
Figure 5 : Comparison of the latency speed plots for ACI-BENCH (top left), MedConv EN (top right), DA (bottom left) and DE (bottom right).
Dimension
Agreement
Pairwise agreement
Krippendorff’s α
Gwet’s AC1
Overall
207/240 (86.3%)
0.767
0.020
0.694
Accurate
29/30 (96.7%)
0.933
0.000
0.929
Thorough
27/30 (90.0%)
0.867
0.284
0.837
Useful
26/30 (86.7%)
0.867
0.442
0.827
Organized
24/30 (80.0%)
0.600
-0.208
0.412
Comprehensible
26/30 (86.7%)
0.733
-0.115
0.653
Appendix
Table 10 : Inter-clinician agreement with the LLM judge across PDSQI dimensions.
Figure 7 : Validation tool showing LLM judgements and dimensions to validate
Figure 8 : Flagged Statements, partially or fully ungrounded, by Dataset x Severity
Sev.
Flagged statement
Category and justification
Corti
S1 test1-D2N108
Rx clindamycin 400 mg PO twice daily for 7 days.
Added route / frequency. The clinician said “you take that,” implying oral administration; adding “PO” would not mislead a reader.
S2 test1-D2N107
Rx physical therapy with home stretches and exercises.
Specificity overreach. The source supports PT with stretches and exercises but does not specify “home.” Clinically implied and unlikely to change management.
S3 test1-D2N109
Air splint applied. Crutches provided. Non-weightbearing until radiograph result.
Contingent to definitive. The clinician describes a plan (“I’m gonna put you in an air splint”), but the note records it as completed, misrepresenting what was done during the visit.
S4 valid-D2N077
Ulnar styloid fracture present.
Fabrication from silence The source mentions ulnar styloid fracture only in a general description of Colles’ fractures, not as a finding for this patient. The note asserts a fracture that was not identified.
Tandem Health
Appendix
Table 11 : Examples of ungrounded statements by system and severity (S1 negligible, S2 minor, S3 moderate, S4 severe). All examples are drawn from ACI-Bench; encounter IDs omit the aci-bench- prefix. Justifications are abridged from the LLM classifier output.
Automatic SOAP note generation can ease the documentation burden on physicians, but existing reasoning methods often omit clinically important information and generate unsupported content. Progress in Thai is further hindered by the lack of publicly available datasets. We propose ASCRIBE, a physician-inspired reasoning framework that ascribes a clinical-significance level to each extracted atomic fact in the conversation before summarization, making a general-purpose LLM a more reliable scribe. We also release ThaiClinicBench, the first de-identified Thai clinical summarization benchmark of real encounters, together with a synthetic training corpus derived from real clinical notes. As a prompt, ASCRIBE outperforms chain-of-thought prompting on GPT-5.4 and Gemini 3.1 Pro across the physician-aligned LLM-judge metrics and improves on standard prompting by up to 10.3 points on the completeness LLM-judge metric. As a GRPO reward, it enables a Gemma-4-E4B model trained solely on synthetic data to match Gemini 3.1 Pro in factual precision and surpass it in completeness. Code and data can be found at https://github.com/loolootech/ascribe.
Large language models (LLMs) can generate or synthesize clinical text for a wide range of applications, from improving clinical documentation to augmenting clinical text analytics. Yet evaluations typically focus on a narrow aspect -- such as similarity or utility comparisons -- even though these aspects are complementary and best viewed in parallel. In this study, we aim to conduct a systematic evaluation of LLM-generated clinical text, which includes intrinsic, extrinsic, and factuality evaluations of synthetic clinical notes rephrased from MIMIC databases at million-note scale. Our analysis demonstrates that synthetic notes preserve core clinical information and predictive utility for coarse-grained tasks despite substantial linguistic changes, but lose fine-grained details for task like ICD coding. We show this loss of detail can be substantially mitigated by rephrasing notes by chunks rather than by the whole note, but at the cost of reduced factual precision under incomplete context. Through fact-checking and error analysis, we further find that synthesis errors are dominated by misinterpretation of clinical context, alongside temporal confusion, measurement errors, and fabricated claims. Finally, we show that the synthetic notes -- despite their task-agnostic nature -- can effectively augment task-specific training for rare ICD codes.
Jinghui Liu, Sarvesh Soni, Anthony Nguyen
Australian e-Health Research Centre, CSIRO, Australia · National Library of Medicine, National Institutes of Health, USA
Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.