Symb-xMIL: Symbolic Explanations for Multiple Instance Learning in Digital Pathology
Organizations: Berlin Institute for the Foundations of Learning and Data, Berlin, Germany · Machine Learning Group, Technische Universität Berlin, Berlin, Germany · Institute of Pathology, Charité Universitätsmedizin, Berlin, Germany · Berlin Institute of Health at Charité – Universitätsmedizin Berlin, BIH Biomedical Innovation Academy, BIH Charité Digital Clinician Scientist Program, Berlin, Germany · Institute of Pathology, Ludwig Maximilian University of Munich, Munich, Germany · Division of Translational Medical Oncology, DKFZ, Heidelberg, Germany · NCT Heidelberg, Heidelberg, Germany · German Cancer Consortium (DKTK), partner site Munich, a partnership between DKFZ and Ludwig-Maximilians-Universität München (LMU), Germany · Department of Artificial Intelligence, Korea University, Seoul, Korea · Max-Planck Institute for Informatics, Saarbrücken, Germany · Department of Chemistry, Chemical Physics Theory Group, University of Toronto, Canada · Vector Institute for Artificial Intelligence, Toronto, Canada · Acceleration Consortium, University of Toronto, Canada
Abstract
Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology. Existing methods primarily rely on heatmaps that highlight influential regions but do not explain how evidence from different tissue regions is combined to produce a prediction. This limits interpretability, especially when decisions depend on interactions between tissue features. We introduce Symbolic explainable MIL (Symb-xMIL), a post-hoc explanation framework that quantifies how a MIL model's behavior aligns with human-readable decision rules, expressed as logical relationships (e.g., AND, OR, NOT) between input features. These alignment scores reveal semantic patterns underlying the model's predictions. We evaluate Symb-xMIL on synthetic and real-world histopathology datasets. On synthetic MIL data, Symb-xMIL reliably recovers ground-truth logical rules. In a clinical tumor detection task, the best-aligned rules uncover heterogeneous decision patterns and expose hidden model errors. On an HPV-prediction task on TCGA-HNSCC, a cohort of head and neck cancer, our framework refines patient survival stratification beyond HPV status with potential clinical relevance. Overall, Symb-xMIL extends MIL explainability beyond visual attribution toward structured, rule-based reasoning, enabling more transparent and semantically grounded interpretation of model predictions.