stat.MLJul 31, 2026

A reproducible and extensible framework for benchmarking competing risks survival models

Authors: Begoña B. SierraColin McLeanPeter S. HallSarah Friedrich-WelzCatalina A. Vallejos

Organizations: Cancer Research UK Scotland Centre · Institute of Genetics and Cancer · University of Edinburgh · Mathematical Statistics and AI in Medicine · University of Augsburg

Abstract

A wide range of statistical and machine learning methods have been proposed for survival analysis with competing risks, where the occurrence of one event (i.e., cancer death) precludes the occurrence of other events (i.e., cardiovascular disease death). Despite these methodological advances, their systematic evaluation and adoption are limited by the lack of comprehensive, reproducible and extensible benchmarking frameworks. We developed an open-source benchmarking framework for competing risks models that enables their systematic comparison across multiple datasets under different aspects of performance; calibration, discrimination, overall prediction error and clinical utility. We additionally introduce an extension of SHAP for competing risks, allowing model-agnostic interpretability of covariates contributions over time. All our code is publicly available via GitHub:https://github.com/BBolosSierra/CompRisksBenchmark

Explore similar work

Date pendingcs.LG

The C-index illusion: discrimination without calibration in published survival models

Recent work has argued normatively, on synthetic data, that evaluating survival models by discrimination alone (concordance index) yields systematically misleading model comparisons, because the metric ignores calibration and time-dependent accuracy. Whether this matters for real, published, non-clinical models has not been tested. We reproduce three published survival-ML models across three structurally distinct domains -- hard-drive failure prediction, peer-to-peer credit default, and user disengagement on digital platforms -- validate our instrument against the anchor paper's own synthetic experiment, and test five pre-registered hypotheses under a Holm-corrected family-wise error rate. Three of five reject (though one pre-registered threshold clears by a narrow margin). A model reproducing the published literature's discrimination almost exactly (C = 0.9595 vs. 0.958 reported) fails a formal calibration test at p < 0.001; a broad feature-ablation search finds no single attribute responsible for its discrimination, so the calibration failure is not a trivial shortcut artifact. A lender's estimated default risk is biased upward by roughly two percentage points, growing to nearly four in the riskiest segment, when loan prepayment is treated as non-informative censoring rather than a competing risk. A platform's churn model shows probability estimates that degrade with the horizon even as global discrimination stays within the pre-registered C-index band. A direct test of whether metric choice inverts model preference does not reject, though with limited power given two to three models per domain; the failure mode we document is better characterized as misplaced confidence in a chosen model than as choosing the wrong one. We release a pre-registered evaluation harness with full code and an annotated notebook, so these results can be verified independently and the audit extended.
Rafael da Silva, Danilo Alvares
Jun 8, 2026stat.AP

A Guide to Estimating Conditional Average Treatment Effects in Competing Risks Settings

Conditional average treatment effects (CATEs) are central to treatment decision-making in personalized medicine. In competing risks settings, estimating CATEs from survival data allows for patient-specific assessments of treatment effectiveness for a specific event of interest while properly accounting for alternative event types. This distinction is essential in the presence of comorbidities, where competing causes of death may otherwise confound the therapeutic benefit. Focusing on right-censored survival times with binary treatment, we examine CATEs defined as covariate-conditional differences in the absolute risk for the event of interest at a fixed time. To this end, we study meta-learners which adapt machine learning algorithms for CATE estimation in competing risks scenarios. We systematically compare six meta-learners, combining Cox regression or random survival forests for risk modeling with elastic net regression or random forests for direct CATE modeling. To provide practical guidance on model selection, we evaluate their performance in multiple simulation settings, that differ in hazard complexity, treatment heterogeneity, treatment assignment, event type distribution and censoring. To facilitate applied use, we provide the R package, crsurvlearners, which implements all considered approaches.
Daniel Klippert, Sarah Friedrich, Markus Pauly
Feb 26, 2025stat.ML

Overcoming Dependent Censoring in the Evaluation of Survival Models

Dependent censoring occurs when the event time and censoring time are not conditionally independent given the observed covariates. This complicates survival model evaluation because widely used metrics, such as the Brier score, typically handle right-censoring using inverse probability of censoring weighting (IPCW). Unfortunately, IPCW is valid only when the estimated censoring distribution is independent of the event time. We propose a dependent Brier score based on an Archimedean copula and the Copula-Graphic estimator, and establish consistency and asymptotic normality of its margin-time estimator. To evaluate the metric, we introduce a semi-synthetic framework that creates realistic dependent censoring while preserving the original covariate structure and known event times. Across 12 datasets, the proposed metric reduces estimation error by 12-16% on average relative to IPCW. Source code is available at https://github.com/thecml/DependentEVAL.
Christian Marius Lillelund, Shi-ang Qi, Russell Greiner