cs.LGOct 1, 2026

Fixing a Model That Learned Worse Cancer Means Lower Risk: Monotonic Constraints in Bladder Cancer Recurrence Prediction

Authors: Saram Abbas, David Thomas, Naeem Soomro, Rishad Shafik, Rakesh Heer, Kabita Adhikari

Organizations: School of Engineering, Newcastle University, Newcastle upon Tyne, UK · Electronics and Computer Science, University of Southampton, Southampton, UK · Department of Urology, Freeman Hospital, Newcastle upon Tyne, UK · Division of Surgery, Imperial College London, London, UK

Abstract

Background and Objective: Clinicians expect recurrence risk to climb with cancer severity. In a UK multicentre trial, an unconstrained XGBoost model learnt that higher tumour stage and carcinoma in situ predicted lower recurrence risk, and discrimination, calibration, and SHAP were all blind to it. We developed a counterfactual testing framework to detect this inversion and a monotonic-constraint framework to remove it without hurting performance. Methods: BOXIT enrolled 472 patients with protocol-mandated cystoscopy across 51 UK sites (2007-2012); 435 had at least two years' follow-up (153 recurrences, 35.2%). We developed a counterfactual direction test and a monotonic-constraint correction, with constraint directions drawn from the EORTC and EAU risk systems, and evaluated both against unconstrained XGBoost and logistic regression on 18 predictors (seven directed) over 50 cross-validation folds. The test worsened each patient on one directed feature at a time to check whether risk fell; SHAP direction and calibration were also assessed. Key Findings and Limitations: Tumour stage and carcinoma in situ were associated with lower recurrence, opposite to medical intuition; the unconstrained model reversed carcinoma in situ counterfactuals in 90.2% of cases and stage in 74.3%. Discrimination (ΔΔAUC 0.005, p=0.47), calibration, and SHAP magnitude were all blind to the inversion. Monotonic constraints eliminated every violation at no cost to discrimination (0.723 vs 0.718) and outperformed EORTC (p=8.9e-16). Limitations: single trial, internal-external validation only. Conclusions and Clinical Implications: A model that had learned this inversion passed every conventional check. A counterfactual direction test, run as a single refit with pre-specified monotonic constraints, catches this failure at no cost to performance and should be routine before clinical deployment.

Figures & tables

Explore similar work

Jul 23, 2026stat.ME

A Multi-Cohort Validation of Censoring-Aware Conformal Lower Predictive Bounds for Pathology Survival Models

Whole-slide survival models commonly provide risk rankings without calibrated statements about individual event times. We evaluate fixed-cutoff drcosarc, a post-hoc conformal wrapper for discrete-time multiple-instance learning survival heads using frozen UNI2-h representations, in an internal 18-configuration sweep across five TCGA cohorts and an external five-configuration evaluation across three CPTAC cohorts. We distinguish configuration--fold--split summaries of the inverse-probability-of-censoring-weighted (IPCW) estimate and median lower predictive bound (LPB) from a hierarchy-aware patient-ensemble estimand of the mean drcosarc--naive LPB difference. At α=0.1α=0.1, the drcosarc IPCW estimate was nearest 0.90 in KIRC, LUAD, and STAD. Patient-ensemble drcosarc--naive intervals excluded zero in KIRC, KIRP, STAD, UCEC, and CPTAC-CCRCC, but included zero in internal LUAD, CPTAC-LUAD, CPTAC-UCEC, and the internal LUSC extension. In a 20-replicate low-censoring semi-synthetic setting with known event times, drcosarc empirical coverage was 0.9129 [0.9053, 0.9207]. An exploratory analysis supported a head-error-by-censoring interaction within that data-generating process. In a two-cohort ABMIL sensitivity analysis, increasing the hazard grid to K=16K=16 raised localized marginal IPCW estimates above the prespecified 0.87 threshold and yielded positive paired LPB differences, although worst-group estimates remained below 0.87. Overall, performance was cohort dependent, and its interpretation changed with the patient-level unit, estimand, and censoring assumptions.
Sep 17, 2026cs.AI

FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity

Deep learning models for multi-modal breast cancer diagnosis achieve high predictive accuracy but remain clinically unacceptable without actionable, counterfactual explanations. Attribution-based methods (LIME, SHAP) are categorically inapplicable to this purpose, as they generate no alternative instances and thus cannot be evaluated on counterfactual quality metrics. This investigation provides empirical evidence that FCA-Guided Counterfactual (FCA-CF) framework that uses a Formal Concept Analysis (FCA) concept lattice as a hard structural constraint on counterfactual search, operating over a multi-modal TCGA-BRCA dataset. We benchmark against four genuine counterfactual methods: Wachter-style CF, DiCE, FACE, and NICE, evaluated on 60 benign-predicted TCGA-BRCA instances. The FCA-CF framework achieves Validity = 1.0000 (100% of counterfactuals successfully flip the prediction), Sparsity = 2.37 features changed (best among all valid methods), and Proximity = 0.900 (normalised L2-based, matching NICE as joint best). The classifier achieves Accuracy = 0.980, F1 = 0.976, ROC-AUC = 0.9947. Ablation analysis confirms that the FCA lattice constraint is the primary sparsity driver (removing it increases sparsity by +40%, p < 0.001, Cohen's d = 0.78), while Phase C greedy refinement accounts for the largest individual contribution (+113% sparsity increase when disabled, p < 0.001, d = 5.01). FCA-guided counterfactual generation achieves a clinically important Pareto-dominant outcome; it is simultaneously the sparsest and among the most proximate of all valid methods, with perfect validity. The emergent sparsity property arising from lattice topology rather than numerical penalty terms constitutes a structurally novel contribution to the counterfactual explanation literature.
Sep 8, 2026eess.IV

CHIMERA Challenge Task 2 and 3: Response Subtypes Classification and Progression Survival Prediction in Bladder Cancer Patients using Multimodal Datasets

High-risk non-muscle-invasive bladder cancer (HR-NMIBC) carries substantial risks of recurrence and progression, while current clinical risk stratification remains limited. CHIMERA was established as a multimodal AI challenge to benchmark prediction in HR-NMIBC under standardized evaluation. Task BRS predicts RNA-seq-defined BCG Response Subtypes from histopathology and structured clinicopathological data, whereas Task Progression models time-to-progression using histopathology, structured data, and RNA sequencing. A multimodal dataset of 368 patients was divided into public training and hidden validation and test sets. In total, 159 submissions were made, and 13 top-performing models were selected for benchmarking. The best models achieved a weighted F1 score of 0.73 for Task BRS and a C-index of 0.68 for Task Progression. Post-challenge analyses revealed task-dependent modality contributions, cohort-dependent performance degradation, and sensitivity to missing structured data. In Task BRS, histopathology partly compensated for pathology-derived structured variables, whereas progression models showed greater dependence on complementary inputs. Cross-model error analysis further identified patients that were consistently difficult across different architectures, with T1 substage associated with prediction difficulty. These findings highlight barriers to transportability and the importance of missingness-aware modeling and independent multi-institutional validation. CHIMERA provides a standardized multimodal benchmark for bladder cancer and a framework for studying not only model performance, but also robustness, information sufficiency, and patient-level prediction failure.