An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer
Abstract
Complex lung cancer decisions can involve several defensible pathways whose eligibility, sequencing and safety depend on unresolved information. Effective support must make explicit how patient conditions govern pathway eligibility, deferral and redirection. MedGPT Clinical Explorer (MCE) organizes alternatives, decision-changing unknowns, safety constraints and fallback into a conditional strategy for clinician review. To evaluate this representation in physician-authored strategies, multidisciplinary experts established case-specific references for 40 cases within a purposive 100-case corpus, and 250 physicians from 98 institutions produced 2,250 strategies under unaided, retrieval-reference and MCE-assisted conditions. MCE-assisted strategies expressed more applicable clinical requirements, measured by the Admissible Pathway Attainment Score (APAS; 0-100), than unaided strategies (adjusted difference, 12.87; 95% CI, 11.18-14.55) and retrieval-reference strategies (5.22; 3.52-6.93). With the same knowledge base available in the retrieval-reference and MCE-assisted conditions, the additional content centered on candidate pathways, decision-critical information and safety constraints. Physicians' whole-strategy acceptability judgments correlated with APAS (Spearman's rho = 0.671), while a complementary relationship audit assessed whether candidates, conditions and subsequent actions were coherently connected. Together, these findings identify two complementary dimensions of open-ended decision support: coverage of clinically relevant content and coherent links among pathways, conditions and subsequent actions. MCE provides a shared decision object that makes consequential omissions and pathway contingencies visible before action; prospective studies should evaluate its effects on clinical workflow and patient outcomes.
Figures & tables
| Characteristic | Overall (N = 100) | Expert-reference set (N = 40) | Source-masked content-evaluation set (N = 60) |
|---|---|---|---|
| Case source | |||
| Internal case | 47 (47.0%) | 40 (100.0%) | 7 (11.7%) |
| Published case | 53 (53.0%) | 0 | 53 (88.3%) |
| Demographic information | |||
| Age, median (IQR), years | 64 (57-69); n = 99 | 65.5 (58.5-69.5); n = 40 | 63 (56.5-68); n = 59 |
| Male | 51 (51.0%) | 28 (70.0%) | 23 (38.3%) |
Appendix figures & tables48 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Expert or review scale | Unit | Prespecified handling rule | Resulting study object |
|---|---|---|---|---|
| Panel composition | 20 multidisciplinary experts | 19 attended the first in-person round; one joined later online rounds | Subsequent case review was undertaken by 14 treatment-focused experts | EASS expert panel |
| Round 1 | 19 experts | Clinical problems, goals, information gaps, pathways, safety and governance in 40 cases | Independent review organized into M1–M5 | Structured case-review material |
| Round 2 | 14 treatment-focused experts | 22 cases reviewed by 4 experts and 18 by 3; 142 expert–case units | Strict majority: at least 2 votes in 3-person review and 3 in 4-person review | Baseline item judgments and items requiring clarification |
| Round 3 | 3 experts per item | 116 independent items; 348 valid responses | Strict majority of valid responses; conservative handling of 1:1:1 splits in information or governance items | Conditions, disagreement and fallback rules |
| M4 safety-rule formation | 40 final rules; 30 safety items carried forward round-2 expert opinions | Round-2 case cards, safety-boundary table and safety review; round-3 pathway-applicability conditions added | Incorporated into APAS rules and linked CSPR gates | Final M4 safety rules |
| Final reference | 240 case-level rules | 40 cases 6 rules | 139 confirmed-acceptable and 101 conditionally acceptable rules | EASS case-specific clinical reference |
| Field | Meaning | Analytic use |
|---|---|---|
| rule_id | Unique within-case rule identifier | Scoring, audit and result linkage |
| case_id | Case identifier | Case-level clustering and mapping |
| module | M1–M5 clinical content domain | Domain summaries |
| final_state | Confirmed- or conditionally acceptable classification | Rule interpretation boundary |
| final_requirement | Clinical requirement to be met | Atomic rule scoring |
| applicability | Conditions under which the rule applies | Denominator determination |
| Item | Definition |
|---|---|
| Score 2 | Fully satisfies the case-level clinical requirement |
| Score 1 | Partly satisfies the requirement or follows a compliant action under its specified condition |
| Score 0 | Omits or conflicts with the requirement, or uses a restricted pathway when its condition is unmet |
| Not applicable | Excluded from APAS numerator and denominator |
| Correct fallback | Excluded from APAS numerator and denominator |
| Reasonable disagreement | Excluded from numerator and denominator when within the prespecified admissible strategy space |
| Gate | Passing requirement |
|---|---|
| CSPR-01 | Strategy addresses the principal clinical problem |
| CSPR-02 | Includes an acceptable or conditionally acceptable main pathway, or correctly chooses verification before commitment |
| CSPR-03 | Does not violate a prespecified high-priority safety constraint in EASS |
| CSPR-04 | Addresses all critical judgment items |
| CSPR-05 | Avoids an unjustified definitive commitment when blocking information is missing |
| CSPR-06 | Respects applicable constraints from comorbidity, organ function, performance status, resources and patient preferences |
| Object | Case input | Verified model or method | Knowledge or retrieval relation | Output | Evaluation use |
|---|---|---|---|---|---|
| MCE | Fixed case input | MedGPT base; state/goal representation, pathway and evidence organization, critique and repair, uncertainty handling and report assembly | Prespecified evidence scope; EASS and CCRR excluded from generation | Strategy Review Pack | M condition, complete MCE in six-configuration comparison, source-masked candidate, and case-condition evaluation |
| U condition | Fixed case input | No additional generation | No additional knowledge material | Physician-submitted strategy | Primary U/R/M evaluation |
| R condition | Fixed case input | GPT-5.4 with LightRAG | Same knowledge base as MCE; does not imply the same retriever, prompt or orchestration | Case-specific reference material | Primary U/R/M evaluation |
| M condition | Fixed case input plus Strategy Review Pack | MCE output available for physician review | Physician could revise and supplement | Physician-submitted strategy | Primary U/R/M evaluation |
| Direct | Fixed case material | Direct generation comparator | No MCE orchestration | Candidate strategy | Secondary source-masked comparison |
| Cross-model other-model strategy output set | 40 expert-reference cases | One strategy from each of GPT-5.4, Claude Opus 4.7 and Gemini 3.1 Pro | No MCE orchestration; not generated in the same batch as complete MCE | Three other-model strategies per case | Supportive system-output comparison in Figure 4a,b |
| Item | Design or observation | Interpretation |
|---|---|---|
| Physicians analyzed | 250 | All completed 9 tasks |
| Complete responses | 2,250 | 250 9 |
| Cases per physician | 9 | No repeated case within an assignment list |
| Conditions per physician | 3 each under U, R and M | Balanced number of conditions within physician |
| Cases | 40/40 covered | Fixed empirical target set |
| Case–condition slots | 120/120 covered | 40 3 |
| Evaluation object | Rater | Repeats or sample | Main output | Inferential role |
|---|---|---|---|---|
| U/R/M physician responses | 3 named AI judges | 3 runs per judge; 9 ratings per response | APAS, M1–M5 and nine CSPR gates | Fixed-panel primary scoring; repeats did not increase the clinical sample |
| Format-aligned MCE output | 3 named AI judges | 3 runs per judge; 9 ratings per output | APAS and M1–M5 | Same repeat and aggregation scheme as physician responses; case-level paired comparison |
| Cross-model other-model strategy outputs | 3 named AI judges | One rating per judge–output; 480 judge–output units | Judge-specific APAS and CSPR | Judges reported separately; no voting, averaging or nine-run aggregation |
| Structural-relation audit of complete MCE outputs | 3 named AI judges | One assessment per judge–output; 120 judge–output units | Five relations, overall relationship closure and defect labels | Judges reported separately; verification tasks nested within outputs |
| Six-configuration functional-module ablation analysis | 3 named AI judges | 40 cases 6 configurations; one report per case–configuration | GPT-5.5 APAS, CSPR and within-case contrasts; rankings from 3 judges | No nine-run aggregation; other judges not pooled with GPT-5.5 at case level |
| AI evaluation of case-condition changes | 3 named AI judges | 48 scenarios and 288 CCRRs; one judgment per judge–requirement | Five-state consensus; requirement and scenario outcomes | Strict-majority consensus; no nine-run aggregation |
| Item | First round | Independent re-review |
|---|---|---|
| Cases | 60 | 60 |
| Modules | 120 | 60 |
| Experts | 6 | A separate group of 6 EASS panelists who did not participate in the first round |
| Per expert | 10 cases, 20 modules | 10 modules |
| MCE versus MDT-Debate | 60 | 30 |
| MCE versus IR-RAG | 20 | 10 |
| Evaluation set | Modules | MCE preferred | No difference | Comparator preferred |
|---|---|---|---|---|
| All first-round modules | 120 | 100 (83.3%) | 6 (5.0%) | 14 (11.7%) |
| First-round modules selected for re-review | 60 | 54 (90.0%) | 2 (3.3%) | 4 (6.7%) |
| Independent re-review | 60 | 57 (95.0%) | 1 (1.7%) | 2 (3.3%) |
| Dimension | Exact agreement across rounds | Exact agreement rate | Wilson 95% CI |
|---|---|---|---|
| D1 key clinical conflicts and treatment goals | 40/60 | 66.7% | 54.1%–77.3% |
| D2 patient state and treatment fit | 44/60 | 73.3% | 61.0%–82.9% |
| D3 safety constraints and unacceptable risk | 48/60 | 80.0% | 68.2%–88.2% |
| D4 decision-changing verification steps | 56/60 | 93.3% | 84.1%–97.4% |
| D5 subsequent pathways and alternatives | 48/60 | 80.0% | 68.2%–88.2% |
| Functional block | Figure 2 nodes | Scientific role | Treatment in the six configurations |
|---|---|---|---|
| Multidisciplinary priors | n0 | Supplies multidisciplinary priors before case-specific construction | Removed only in the corresponding grouped ablation |
| Case-state, goal and constraint representation | n1-n2 | Compiles the fixed case and represents goals, constraints and decision-changing unknowns | Retained in the MCE-derived ablation configurations |
| Pathway and qualifying-evidence construction | n3-n5 | Constructs foundational and derived pathways and links evidence that supports, limits or leaves applicability unresolved | Removed jointly with n6-n8 in the n3-n8 ablation |
| Simulation, critique and repair | n6-n8 | Tests candidate benefits, risks, feasibility and dependencies and repairs or removes candidates | Removed alone in one grouped ablation and jointly with n3-n5 in another |
| Candidate reranking and candidate-linked verification | n9 | Reranks candidate strategies and forms verification objects linked to subsequent actions | Candidate-strategy reranking was retained in both stage-removal configurations; complete retention of every verification behavior was not assumed |
| Stopping assessment and report assembly | n10 | Determines whether the object returns to an earlier stage or is assembled as a Strategy Review Pack | Retained in the MCE-derived ablation configurations |
| Configuration | Multidisciplinary priors | Pathways and qualifying evidence | Simulation, critique and repair | Candidate-strategy reranking | Retrieval augmentation | Remaining MCE orchestration |
|---|---|---|---|---|---|---|
| Complete MCE | Retained | Retained | Retained | Retained | Retained | Retained |
| Without multidisciplinary priors | Removed | Retained | Retained | Retained | Retained | Retained |
| Without pathway/evidence and review stages | Retained | Removed | Removed | Retained | Configuration-limited | Other stages retained |
| Without simulation, critique and repair | Retained | Retained | Removed | Retained | Retained | Other stages retained |
| MedGPT base direct generation | Not applicable | Not applicable | Not applicable | Not applicable | Removed | Removed |
| MedGPT-LightRAG | Not applicable | Not applicable | Not applicable | Not applicable | Retained; same knowledge base as MCE | Removed |
| Configuration | Cases | Mean APAS (95% CI) | CSPR passing/40 (%) |
|---|---|---|---|
| Complete MCE | 40 | 91.91 (89.12–94.71) | 37/40 (92.5%) |
| Without multidisciplinary priors | 40 | 75.29 (70.69–79.89) | 32/40 (80.0%) |
| Without pathway/evidence and review stages | 40 | 84.19 (80.94–87.44) | 37/40 (92.5%) |
| Without simulation, critique and repair | 40 | 88.31 (85.42–91.20) | 37/40 (92.5%) |
| MedGPT base direct generation | 40 | 68.09 (60.57–75.61) | 25/40 (62.5%) |
| MedGPT-LightRAG | 40 | 77.57 (72.13–83.01) | 33/40 (82.5%) |
| Comparison | Mean APAS difference (95% CI) |
|---|---|
| Complete MCE versus without multidisciplinary priors | 16.62 (11.84–21.40) |
| Complete MCE versus without pathway/evidence and review stages | 7.72 (4.87–10.57) |
| Complete MCE versus without simulation, critique and repair | 3.60 (1.72–5.49) |
| Complete MCE versus MedGPT base direct generation | 23.82 (15.95–31.70) |
| Complete MCE versus MedGPT-LightRAG | 14.34 (9.11–19.57) |
| Judge pair | Spearman’s | Ranking units |
|---|---|---|
| GPT-5.5 versus Gemini 3.1 Pro Preview | 1.000 | 6 configurations |
| GPT-5.5 versus Claude Opus 4.7 | 0.943 | 6 configurations |
| Gemini 3.1 Pro Preview versus Claude Opus 4.7 | 0.943 | 6 configurations |
| State | Definition | Main denominator handling |
|---|---|---|
| Pass | Output meets the prespecified requirement | Included in pass/clear-gap denominator |
| Clear gap | Output clearly omits or violates the requirement | Included in pass/clear-gap denominator |
| Clinical review | Structured judging alone does not support a reliable determination | Counted as not passing in the conservative descriptive denominator; separate in the binary-evaluable analysis |
| Not applicable | Requirement does not apply to the scenario | Reported separately; excluded from both proportion denominators |
| Not assessable | Material or returned output is insufficient for assessment | Counted as not passing in the conservative descriptive denominator; separate in the binary-evaluable analysis |
| Stage | Entry rule | Rater or reviewer | Interpretation |
|---|---|---|---|
| AI consensus | Strict majority of three judges; a three-state split assigned clinical review | 3 named AI judges | Structured localization of potential gaps, not a clinical reference standard |
| Review cohort | Clear gap, clinical review, judge disagreement, missing judge result, or prespecified high-priority or ungraded safety requirement | 2 physicians independently | Did not define a human reference standard |
| Third-physician escalation | All safety requirements and all physician disagreements | Third physician with access to earlier opinions | Targeted escalation, not blinded independent scoring |
| Complete scenario closure | All 6 requirements returned; at least one entered the main denominator; all in-denominator requirements passed; no clear gap, clinical review or not-assessable state | Derived separately for each rating source | Not a clinical safety or efficacy endpoint |
| Unit | State or analysis | Count | Denominator and interpretation |
|---|---|---|---|
| Requirement | Pass | 228 | Consensus among 288 prespecified requirements |
| Requirement | Clinical review | 45 | Counted as not passing in the conservative descriptive analysis; separate in binary-evaluable analysis |
| Requirement | Clear gap | 14 | Together with pass, formed 242 binary-evaluable requirements |
| Requirement | Not applicable | 1 | Excluded from denominators 242 and 287 |
| Requirement | Not assessable | 0 | Counted as not passing in conservative analysis; none observed |
| Requirement | Conservative descriptive pass proportion | 228/287 (79.4%) | Clinical review and not assessable counted as not passing; not applicable excluded |
| Condition category | Scenarios | Complete closure | Clinical review | Clear gap |
|---|---|---|---|---|
| All scenarios | 48 | 13 | 23 | 12 |
| BASE baseline anchor | 16 | 5 | 8 | 3 |
| SAFE safety constraint | 11 | 4 | 5 | 2 |
| INFO decision-critical information | 5 | 1 | 2 | 2 |
| SEQ pathway sequence | 8 | 3 | 4 | 1 |
| GOV reassessment and governance | 8 | 0 | 4 | 4 |
| Mechanism | Mechanism labels | Scenario cards | Source templates |
|---|---|---|---|
| Pathway prerequisite unmet | 8 | 8 | 6 |
| Sequence or fallback violation | 6 | 6 | 5 |
| Core clinical problem not covered | 2 | 2 | 1 |
| Decision-critical information unresolved | 2 | 2 | 1 |
| Safety-constraint relationship unmet | 2 | 2 | 1 |
| Comparison | Level | Objects | Exact agreement | Gwet AC1 (95% CI) |
|---|---|---|---|---|
| Physician 1 versus physician 2 | Rule | 288 CCRRs; 16 source templates | 67.01% | 0.607 (0.530–0.684) |
| Physician 1 versus physician 2 | Scenario card | 48 scenarios; 16 source templates | 93.75% | 0.892 (0.757–1.000) |
| Physician 1 versus AI consensus | Rule | 288 CCRRs; 16 source templates | 39.93% | 0.301 (0.204–0.402) |
| Physician 2 versus AI consensus | Rule | 288 CCRRs; 16 source templates | 39.93% | 0.303 (0.199–0.399) |
| Physician 1 versus AI consensus | Scenario card | 48 scenarios; 16 source templates | 58.33% | 0.311 ( 0.050 to 0.618) |
| Physician 2 versus AI consensus | Scenario card | 48 scenarios; 16 source templates | 56.25% | 0.244 ( 0.149 to 0.600) |
| Modification category | Pairs | Toward complete closure | Unchanged | Toward clear gap |
|---|---|---|---|---|
| SAFE safety constraint | 11 | 3 | 6 | 2 |
| INFO decision-critical information | 5 | 0 | 3 | 2 |
| SEQ pathway sequence | 8 | 2 | 3 | 3 |
| GOV reassessment and governance | 8 | 1 | 4 | 3 |
| Safety label | Requirements | Pass | Clinical review | Clear gap |
|---|---|---|---|---|
| P0 | 9 | 7 | 0 | 2 |
| P1 | 24 | 23 | 1 | 0 |
| Ungraded safety condition | 12 | 10 | 2 | 0 |
| No case-specific P0/P1 identified | 3 | 3 | 0 | 0 |
| Analysis | Model or method | Unit or clustering | Diagnostics and fallback | Multiplicity or interval |
|---|---|---|---|---|
| Primary APAS | APAS ~ condition case + task position + (1 | physician) | Physician–case–condition response | REML; singularity tolerance 1e-4; convergence and warnings recorded | Holm adjustment for M–U and M–R P values; Wald 95% CI |
| APAS standardization | Equally weighted 40-case 9-position reference grid | Fixed-case empirical target set | Physician random effect set to 0 | Delta-method standard error |
| CSPR | CSPR ~ condition + case + task position + (1 | physician) ; fixed effects transformed with expit and averaged over the 40 9 grid | 2,237 evaluable responses | Common-condition-effect model used after instability of the initial case-by-condition binomial model; available diagnostics did not identify a more specific trigger | Probability differences and delta-method standard errors; Holm adjustment for M–U and M–R P values |
| Judge stability | Nine judge-run models, equal-weight panel and leave-one-judge-out panels | Same 2,250 responses | Repeated ratings were not independent samples | Direction and interval reported separately |
| Professional-title-group analysis | Condition case + condition professional-title group + task position + physician random effect | Physician–case response | Joint four-degree-of-freedom Wald test | Absence of detected interaction was not interpreted as equivalence |
| M1–M5 domains | Case-heterogeneous models with the APAS structure | Responses with an applicable domain score | Non-applicable records not imputed | Holm adjustment across ten contrasts; interactions post hoc |
| Domain | Contrast | APAS difference | 95% CI | Holm-adjusted P value | Evaluable responses by condition |
|---|---|---|---|---|---|
| M1 clinical problem and treatment goals | M–U | 13.10 | 11.07–15.12 | U 750; R 750; M 750 | |
| M1 clinical problem and treatment goals | M–R | 4.37 | 2.32–6.42 | U 750; R 750; M 750 | |
| M2 decision-critical information | M–U | 8.59 | 6.35–10.83 | U 429; R 465; M 505 | |
| M2 decision-critical information | M–R | 4.88 | 2.61–7.15 | U 429; R 465; M 505 | |
| M3 candidate clinical pathways | M–U | 20.64 | 18.15–23.13 | U 742; R 731; M 735 | |
| M3 candidate clinical pathways | M–R | 8.59 | 6.07–11.12 | U 742; R 731; M 735 |
| Professional-title stratum | Physicians | Contrast | APAS difference | 95% CI |
|---|---|---|---|---|
| Senior | 68 | M–U | 14.61 | 11.36–17.86 |
| Senior | 68 | M–R | 5.16 | 1.89–8.43 |
| Intermediate | 91 | M–U | 11.68 | 8.89–14.48 |
| Intermediate | 91 | M–R | 4.24 | 1.43–7.04 |
| Junior | 91 | M–U | 12.76 | 9.97–15.54 |
| Junior | 91 | M–R | 6.22 | 3.42–9.03 |
| Output | AI judge | Mean APAS | 95% CI | CSPR passing/40 | CSPR pass rate |
|---|---|---|---|---|---|
| Complete MCE | Gemini 3.1 Pro Preview | 96.23 | 93.64–98.82 | 36/40 | 90.0% |
| Complete MCE | GPT-5.5 | 94.71 | 91.35–98.07 | 36/40 | 90.0% |
| Complete MCE | Claude Opus 4.7 | 97.72 | 95.94–99.50 | 39/40 | 97.5% |
| GPT-5.4 | Gemini 3.1 Pro Preview | 71.06 | 62.33–79.78 | 29/40 | 72.5% |
| GPT-5.4 | GPT-5.5 | 68.75 | 61.30–76.20 | 27/40 | 67.5% |
| GPT-5.4 | Claude Opus 4.7 | 74.34 | 67.40–81.28 | 35/40 | 87.5% |
| Output | AI judge | M1 | M2 | M3 | M4 | M5 |
|---|---|---|---|---|---|---|
| Complete MCE | Gemini 3.1 Pro Preview | 96.25 (92.51–99.99) | 96.05 (90.35–101.76) † | 93.75 (87.49–100.01) | 97.50 (92.60–102.40) | 98.75 (96.30–101.20) |
| Complete MCE | GPT-5.5 | 94.38 (89.91–98.84) | 96.25 (90.83–101.67) | 91.25 (83.49–99.01) | 96.25 (90.83–101.67) | 96.25 (92.12–100.38) |
| Complete MCE | Claude Opus 4.7 | 99.38 (98.15–100.60) | 97.50 (94.08–100.92) | 95.00 (89.13–100.87) | 97.50 (94.08–100.92) | 97.50 (94.08–100.92) |
| GPT-5.4 | Gemini 3.1 Pro Preview | 74.38 (64.53–84.22) | 76.32 (64.78–87.85) † | 53.75 (42.44–65.06) | 72.50 (60.89–84.11) | 77.50 (68.25–86.75) |
| GPT-5.4 | GPT-5.5 | 78.12 (68.65–87.60) | 70.00 (59.59–80.41) | 52.50 (40.89–64.11) | 70.00 (59.59–80.41) | 61.25 (53.82–68.68) |
| GPT-5.4 | Claude Opus 4.7 | 83.12 (75.20–91.05) | 71.25 (62.74–79.76) | 66.25 (55.50–77.00) | 72.50 (63.94–81.06) | 67.50 (59.23–75.77) |
| Comparison | AI judge | M1 | M2 | M3 | M4 | M5 |
|---|---|---|---|---|---|---|
| Complete MCE minus GPT-5.4 | Gemini 3.1 Pro Preview | 21.88 (11.77–31.98) | 19.74 (6.65–32.83) † | 40.00 (25.45–54.55) | 25.00 (11.87–38.13) | 21.25 (12.04–30.46) |
| Complete MCE minus GPT-5.4 | GPT-5.5 | 16.25 (5.64–26.86) | 26.25 (15.73–36.77) | 38.75 (24.94–52.56) | 26.25 (15.73–36.77) | 35.00 (27.00–43.00) |
| Complete MCE minus GPT-5.4 | Claude Opus 4.7 | 16.25 (8.30–24.20) | 26.25 (17.66–34.84) | 28.75 (17.72–39.78) | 25.00 (16.40–33.60) | 30.00 (21.55–38.45) |
| Complete MCE minus Claude Opus 4.7 | Gemini 3.1 Pro Preview | 33.75 (22.58–44.92) | 31.58 (16.15–47.01) † | 50.00 (35.96–64.04) | 26.25 (12.67–39.83) | 31.25 (19.78–42.72) |
| Complete MCE minus Claude Opus 4.7 | GPT-5.5 | 26.88 (15.45–38.30) | 40.00 (27.25–52.75) | 52.50 (37.22–67.78) | 41.25 (28.17–54.33) | 40.00 (27.25–52.75) |
| Complete MCE minus Claude Opus 4.7 | Claude Opus 4.7 | 27.50 (18.94–36.06) | 32.50 (22.86–42.14) | 38.75 (26.35–51.15) | 31.25 (22.18–40.32) | 25.00 (15.08–34.92) |
| Outcome or domain | Format-aligned MCE mean | Physician case mean | Paired difference (format-aligned MCE minus physician response) | Case-cluster 95% CI | Paired cases |
|---|---|---|---|---|---|
| APAS versus U | 85.29 | 53.26 | 32.03 | 26.48–37.09 | 40 |
| APAS versus R | 85.29 | 61.00 | 24.29 | 19.94–28.60 | 40 |
| APAS versus M | 85.29 | 66.18 | 19.11 | 15.58–22.33 | 40 |
| Non-whitespace characters versus M | 363.6 | 611.1 | 247.5 | 313.4 to 185.2 | 40 |
| M1 clinical problem and treatment goals versus M | 84.58 | 63.95 | 20.63 | 16.48–24.51 | 40 |
| M2 decision-critical information versus M | 91.39 | 81.66 | 9.73 | 5.75–13.59 | 40 |
| Estimand | Component clinical-event record (n=12) | Follow-up without a recorded component event (n=19) | Between-group difference, event-record minus no-record group (95% CI) |
|---|---|---|---|
| APAS under U | 55.27 (52.51–58.03) | 49.57 (47.09–52.05) | 5.70 (2.84–8.57) |
| APAS under R | 64.44 (61.52–67.35) | 56.26 (53.80–58.73) | 8.17 (5.18–11.17) |
| APAS under M | 68.47 (65.62–71.31) | 64.37 (61.93–66.81) | 4.10 (1.18–7.01) |
| M–U | 13.20 (10.03–16.37) | 14.80 (12.29–17.31) | 1.61 ( 5.74 to 2.53) |
| M–R | 4.03 (0.74–7.33) | 8.11 (5.62–10.60) | 4.08 ( 8.28 to 0.12) |
| Domain | Event-record group M–U (n=14; 95% CI) | No-record group M–U (n=20; 95% CI) | Event-record group M–R (n=14; 95% CI) | No-record group M–R (n=20; 95% CI) |
|---|---|---|---|---|
| M1 clinical problem and treatment goals | 8.89 (5.36–12.43) | 15.86 (12.90–18.82) | 2.08 ( 1.56 to 5.71) | 6.93 (4.01–9.85) |
| M2 decision-critical information | 8.65 (4.76–12.54) | 9.65 (6.39–12.91) | 2.32 ( 1.68 to 6.32) | 8.03 (4.80–11.25) |
| M3 candidate clinical pathways | 18.91 (14.56–23.26) | 20.09 (16.45–23.74) | 5.66 (1.18–10.13) | 11.47 (7.87–15.06) |
| M4 safety boundaries and risk governance | 10.69 (7.16–14.22) | 13.11 (10.16–16.07) | 4.62 (1.00–8.25) | 5.38 (2.46–8.29) |
| M5 reassessment and governance | 3.25 (0.38–6.12) | 5.44 (3.04–7.85) | 3.59 (0.64–6.55) | 4.63 (2.26–7.00) |
| Analysis | Groups and case counts | APAS under M | M–U | M–R | Between-group difference in contrast (95% CI) |
|---|---|---|---|---|---|
| All records with available follow-up | Component event record, 14; follow-up without recorded component event, 20 | 66.71 vs 64.46 | 10.69 vs 14.20 | 3.60 vs 7.45 | M–U: 3.50 ( 7.44 to 0.43); M–R: 3.85 ( 7.81 to 0.10) |
| Recorded-death stratum | Death recorded, 5; follow-up without recorded death, 29 | 69.01 vs 64.76 | 5.94 vs 13.93 | 0.89 vs 6.72 | M–U: 7.99 ( 13.57 to 2.41); M–R: 5.83 ( 11.42 to 0.24) |
| Recorded-death stratum excluding the record with a reversed date sequence | Death recorded, 4; follow-up without recorded death, 29 | 67.90 vs 64.76 | 7.42 vs 13.93 | 1.10 vs 6.72 | M–U: 6.51 ( 12.65 to 0.36); M–R: 5.62 ( 11.82 to 0.58) |
| Post hoc stratum | Cases | M–U (95% CI) | M–R (95% CI) |
|---|---|---|---|
| Dual goals | 12 | 7.45 (4.34–10.57) | 1.90 ( 1.26 to 5.07) |
| Primary goal with fallback goal | 27 | 15.05 (12.93–17.16) | 6.60 (4.46–8.73) |
| Diagnosis, staging, tumor origin or biological boundary | 11 | 14.61 (11.37–17.85) | 1.59 ( 1.71 to 4.88) |
| Local-treatment resectability and local control | 7 | 10.57 (6.33–14.80) | 5.68 (1.47–9.89) |
| Systemic-treatment selection, resistance and sequence | 5 | 8.53 (3.71–13.34) | 7.31 (2.10–12.52) |
| Treatment toxicity, comorbidity and tolerance constraints | 16 | 13.68 (10.92–16.44) | 6.80 (4.00–9.60) |
| Subgroup | Cases | M2 M–U | M3 M–U | M4 M–U |
|---|---|---|---|---|
| Relevant record without a recorded death | 7 | 17.01 | 28.04 | 16.96 |
| Estimand | Unit | Result | 95% CI |
|---|---|---|---|
| Reviewer 1 four-category acceptability distribution | 120 outputs; unacceptable / major revision or verification / acceptable with minor revision / acceptable for clinical discussion | 12/34/19/55 | — |
| Reviewer 2 four-category acceptability distribution | Same | 8/31/23/58 | — |
| Reviewer 1 binary acceptable | 120 outputs | 74/120 (61.7%) | — |
| Reviewer 2 binary acceptable | Same | 81/120 (67.5%) | — |
| Acceptable to both reviewers | Same | 71/120 (59.2%) | — |
| Reviewer 1 major clinical defect | Same | 38/120 (31.7%) | — |
| Estimand | Unit | Result | 95% CI or limits |
|---|---|---|---|
| Rule-level exact agreement | 240 rules; 40 complete MCE reports | 188/240 (78.33%) | — |
| Rule-level adjacent agreement | Same | 238/240 (99.17%) | — |
| Linearly weighted Gwet AC2 | Same | 0.842 | 0.785–0.893 |
| Linearly weighted Cohen | Same | 0.411 | 0.257–0.558 |
| Case-level ICC(A,1) | 40 paired reports | 0.694 | 0.468–0.835 |
| Case-level Lin CCC | Same | 0.689 | — |
| Variant type | Triplets | Overall APAS difference, variant minus original (95% CI) | Decrease | No change | Increase |
|---|---|---|---|---|---|
| Prespecified section-order variant | 36 | +0.85 ( 0.44 to 2.31) | 15/36 | 2/36 | 19/36 |
| Rule-linked target-content-deletion variant | 36 | 0.66 ( 1.69 to 0.36) | 20/36 | 1/36 | 15/36 |
| Target domain | Triplets | Target-domain APAS difference, variant minus original (95% CI) |
|---|---|---|
| M2 decision-critical information | 9 | 0.32 ( 3.41 to 2.13) |
| M3 candidate clinical pathways | 9 | 3.70 ( 8.02 to 0.00) |
| M4 safety boundaries and risk governance | 9 | 0.46 ( 3.86 to 4.32) |
| M5 reassessment and governance | 9 | 1.85 ( 4.94 to 1.23) |
| Variant type | Intended transformation | Triplets | Main comparison |
|---|---|---|---|
| Prespecified section-order variant | Change section order while retaining strategy content | 36 | Overall APAS, variant minus original |
| Rule-linked target-content-deletion variant | Delete designated clinical content linked to M2, M3, M4 or M5 | 36 (9 per domain) | Overall and target-domain APAS, variant minus original |
| Gray-zone criterion | Operational definition |
|---|---|
| Guideline recommendation unclear | The guideline recommendation is unclear, or the combination of patient features does not map directly to a standard guideline scenario. |
| Major conflict from comorbidity or treatment constraints | An important non-cancer medical constraint creates conflict between otherwise plausible pathways. |
| Documented decision disagreement | The multidisciplinary record contains at least two distinct treatment recommendations or a documented disagreement between senior experts. |
| Pathway sequence or strategy choice may materially affect outcome | Treatment sequence, pathway combination or ordering may materially alter later options or outcomes. |
| Interacting factors limit direct application of a standard pathway | Molecular, pathological, imaging, comorbidity, treatment-history or other factors jointly prevent direct application of a standard pathway. |
| Prespecified attribute (non-exclusive) | Overall (N=100) | Expert-reference set (N=40) | Source-masked content-evaluation set (N=60) |
|---|---|---|---|
| Guideline recommendation unclear | 74 (74.0%) | 26 (65.0%) | 48 (80.0%) |
| Major conflict from comorbidity or treatment constraints | 47 (47.0%) | 26 (65.0%) | 21 (35.0%) |
| Documented decision disagreement | 1 (1.0%) | 0 | 1 (1.7%) |
| Pathway sequence or strategy choice may materially affect outcome | 45 (45.0%) | 1 (2.5%) | 44 (73.3%) |
| Interacting factors limit direct application of a standard pathway | 63 (63.0%) | 11 (27.5%) | 52 (86.7%) |
| At least two prespecified gray-zone criteria | 78 (78.0%) | 24 (60.0%) | 54 (90.0%) |
| Primary clinical-decision scenario | Overall (N=100) | Expert-reference set (N=40) | Source-masked content-evaluation set (N=60) |
|---|---|---|---|
| Diagnosis, staging, tumor origin or biological boundary | 19 (19.0%) | 11 (27.5%) | 8 (13.3%) |
| Local-treatment resectability and local control | 12 (12.0%) | 7 (17.5%) | 5 (8.3%) |
| Systemic-treatment selection, resistance and sequence | 22 (22.0%) | 5 (12.5%) | 17 (28.3%) |
| Treatment toxicity, comorbidity and tolerance constraints | 32 (32.0%) | 16 (40.0%) | 16 (26.7%) |
| Prioritization of coexisting disease or multiple primary tumors | 6 (6.0%) | 1 (2.5%) | 5 (8.3%) |
| Acute critical illness and initial symptom stabilization | 9 (9.0%) | 0 | 9 (15.0%) |
| Judge | D1 candidate-pathway differentiation | D2 action-premise relation | D3 decision-changing unknown | D4 verification-candidate relation | D5 reassessment/fallback relation | All five complete | Overall relationship closure |
|---|---|---|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 40/40 | 37/40 | 40/40 | 38/40 | 40/40 | 36/40 | 36/40 |
| GPT-5.5 | 40/40 | 31/40 | 40/40 | 37/40 | 40/40 | 30/40 | 30/40 |
| Claude Opus 4.7 | 40/40 | 40/40 | 40/40 | 40/40 | 40/40 | 40/40 | 40/40 |
| Judge | Verification tasks | Subsequent action after support | Subsequent action after refutation | Subsequent action if unresolved |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 137 | 137/137 | 137/137 | 137/137 |
| GPT-5.5 | 145 | 145/145 | 145/145 | 144/145 |
| Claude Opus 4.7 | 135 | 135/135 | 135/135 | 135/135 |
| Total judge–task observations | 417 | 417/417 | 417/417 | 416/417 |
| Anonymized output | Gemini 3.1 Pro Preview | GPT-5.5 | Claude Opus 4.7 |
|---|---|---|---|
| Output 1 | CONFLICT | CONFLICT | NONE |
| Output 2 | NONE | CONFLICT | NONE |
| Output 3 | NONE | CONFLICT | NONE |
| Output 4 | NONE | CONFLICT | NONE |
| Output 5 | MISCONNECTION | CONFLICT/MISCONNECTION | NONE |
| Output 6 | NONE | CONFLICT | NONE |
| Review target | Items | Result |
|---|---|---|
| Still-unresolved branches | 10 | Subsequent action confirmed in 9; one omission confirmed |
| Additional judgment rationale | 8 | Text evidence and location recorded for supportive, refuting and unresolved states |
| Total reviewed objects | 18 | All checked against the corresponding output |
| Study characteristic | Description | Supporting documentation |
|---|---|---|
| Cases and study sets | 100 cases: 40 in the expert-reference set and 60 in the source-masked content-evaluation set | Case identifier and study-set membership |
| Case sources | 47 restricted internal clinical cases and 53 published cases | Source type and public-source identifier |
| Inclusion and grouping | Purposively assembled using prespecified gray-zone attributes; 40 internal cases formed the expert-reference set and the remaining 60 cases formed the source-masked content-evaluation set; 0 actively excluded | Analysis-set and source-set labels |
| Fixed case input | Six fields: basic information, chief complaint, present illness, past and other medical history, investigations, and diagnosis; excluded management, treating-physician recommendations, subsequent plans and outcomes | Input structure version and cutoff |
| De-identification and release boundary | Internal cases were limited, de-identified restricted data; submission materials exclude case narratives and access logs | De-identification category and access conditions |
| Identifier and integrity | 100 unique case identifiers and unique decision-input and full-record integrity signatures; no identifier or input overlap between study sets | Case identity and content integrity |