Same Predictions, Different Harms: Causal Auditing of Patient World Models
Organizations: Imperial College London, London, United Kingdom
Abstract
Patient world models used for clinical trial simulation can agree on transition kernels and arm-specific risks, yet disagree on the fraction of patients harmed by switching treatment---the counterfactual quantity that matters for intervention-aware reasoning. We audit this reliability gap in a two-stage shared-response SCM: a categorical intermediate health state is followed by common terminal care. Under independent stages, the sharp harm interval has closed-form endpoints for at most three intermediate states, with an exactness boundary at four states. Declared dependence and response-mismatch budgets yield calibrated outer bounds when stage independence or complete mediation is relaxed; in a symmetric three-state model the entire sensitivity frontier is sharp, , and shows exactly how budgets erase the gain over endpoint-only bounds. Two eight-variable response LPs propagate interventional uncertainty for finite-sample audits. Exact witnesses verify attainability. On public clinical simulators (EpiCare; sepsis), native configurations show little resolved stage dependence and no additional joint-compatibility gain over pairwise transport---honest negative results for reliability claims. All experiments are locally reproducible; guarantees remain conditional on the stated causal model. The results provide a concrete protocol for deciding when a patient world model is safe to trust for counterfactual harm.
Figures & tables
| Model class | Guarantee and remaining restriction |
|---|---|
| Independent shared response, | Sharp endpoints; arbitrary risks. |
| Independent shared response, general | Sharp lower; cut/pair upper generally outer. Exact under a dominant pooled state. |
| Joint dependence and direct-response budgets | General calibrated outer interval; sharp symmetric three-state frontier. |
| Ternary entrance and binary Markov suffix | Sharp at any suffix length; requires independent noises and common care. |
| Unknown interventional primitives | Conservative simultaneous confidence outer interval under valid sampling. |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | LH | Projection | CP | H | 3-event | |
|---|---|---|---|---|---|---|
| 1 | 100 | 0.6353 | 0.6166 | 0.5535 | 0.5746 | 0.5500 |
| 1 | 500 | 0.5600 | 0.4757 | 0.5243 | 0.5341 | 0.5227 |
| 1 | 2000 | 0.4583 | 0.4040 | 0.5121 | 0.5170 | 0.5112 |
| 1 | 10000 | 0.3890 | 0.3644 | 0.5053 | 0.5075 | 0.5049 |
| 2 | 100 | 0.7600 | 0.6818 | 0.5044 | 0.5766 | 0.4882 |
| 2 | 500 | 0.5149 | 0.3986 | 0.3894 | 0.4233 | 0.3839 |
| Setting | ||||
|---|---|---|---|---|
| 1 | ||||
| 2 | ||||
| 3 | ||||
| 4 | ||||
| 5 | ||||
| 6 |
| Setting | ||||
|---|---|---|---|---|
| 1 | 0.635 / 0.621 | 0.560 / 0.553 | 0.458 / 0.527 | 0.389 / 0.512 |
| 2 | 0.760 / 0.726 | 0.515 / 0.495 | 0.352 / 0.396 | 0.248 / 0.342 |
| 3 | 0.652 / 0.637 | 0.564 / 0.557 | 0.467 / 0.501 | 0.361 / 0.447 |
| 4 | 0.680 / 0.666 | 0.589 / 0.589 | 0.448 / 0.517 | 0.342 / 0.462 |
| 5 | 0.537 / 0.522 | 0.435 / 0.454 | 0.317 / 0.427 | 0.252 / 0.412 |
| 6 | 0.426 / 0.411 | 0.337 / 0.330 | 0.299 / 0.296 | 0.279 / 0.277 |
| States | Mixing | Positive gaps | Mean gap | Maximum gap | Max. cut gap |
|---|---|---|---|---|---|
| 3 | 0 | 0/30 | 0 | 0 | 0 |
| 3 | 2/30 | .00157 | .02466 | 0 | |
| 3 | 1 | 30/30 | .06828 | .11347 | 0 |
| 4 | 0 | 0/30 | 0 | 0 | 0 |
| 4 | 9/30 | .00605 | .04682 | .00995 | |
| 4 | 1 | 30/30 | .01665 | .05270 | .01993 |
| Setting | ||||
|---|---|---|---|---|
| 1 | 1.000 | 0.999 | 0.999 | 0.998 |
| 2 | 1.000 | 0.998 | 1.000 | 0.998 |
| 3 | 1.000 | 1.000 | 0.998 | 1.000 |
| 4 | 0.998 | 0.999 | 1.000 | 0.999 |