cs.LGSep 27, 2026

T-MoXAI: A Hierarchical Explainability Framework for Temporal Multimodal Data

Authors: Ali Inha, Mo Vali, Saaliha Vali, Pietro Liò, Meen-Yau Thum

Organizations: Department of Computer Science, University of Cambridge, Cambridge, UK · Cavendish Laboratory, Department of Physics, University of Cambridge, Cambridge, UK · Imperial College Healthcare NHS Trust, London, UK · Lister Fertility Clinic, The Lister Hospital, HCA UK, UK

Abstract

Artificial Intelligence (AI) models for temporal multimodal data have potential in healthcare and agriculture, but their opacity can limit trust and adoption. We introduce T-MoXAI (Temporal Multimodal eXplainable AI), a hierarchical framework explaining (1) when timepoints influence predictions, using temporal Shapley values; (2) which modalities contribute at those moments, using attention analysis; and (3) what features or image regions drive decisions, using gradient based attribution. A transformer based architecture handles irregular temporal sequences and heterogeneous data, generating all three explanation levels in under one second for interactive decision support. We evaluate the framework on two real world tasks: predicting IVF treatment outcomes from ultrasound sequences and clinical measurements (AUC 0.660 despite significant class imbalance), and forecasting wheat yield from temporal RGB imagery and phenotypic traits (R2R^2 0.265 amid substantial environmental variability). Ablation studies indicate that temporal modelling is critical in both domains: removing it reduces performance to the equivalent of random guessing. Temporal ROAR experiments provide evidence that the explanations reflect the model's reasoning process. With a unified, domain agnostic architecture and open source implementation, T-MoXAI provides a baseline for temporal multimodal XAI, addressing fragmentation in the field and supporting applications where understanding decisions is as important as predictive accuracy.

Figures & tables

Explore similar work

Jun 25, 2026cs.LG

Global Explanations for Multivariate Time Series Forecasting Models via KK-Order Markov Approximations

While many explainable AI (XAI) methods have been proposed, most are not designed for time-series forecasting models and often rely on the implicit assumption that timestamp features are independent. This assumption ignores the fundamental property of temporal dependence and can lead to explanations that violate the sequential and causal structure of the data. We introduce \textsc{KARMA}, a method for explaining time-series predictors by constructing a Markov surrogate model that captures the temporal dependencies learned by the predictor. Our approach revolves around three main aspects: identifying the minimal history length KK that is predictively sufficient for the model, estimating the best-fitting KK-order Markov transition kernel from the discretized history space, and a five-level global explanation hierarchy that can be derived from the Markov transition kernel, which we illustrate using real-world weather data (Beijing PM 2.5). We also certify using complex synthetic data with known true causal edges that KARMA (i) recovers the data causal structure as learned by the model via a controlled experiment and (ii) identifies temporal dependencies better than established attribution methods such as TimeSHAP.
Mar 4, 2026cs.LG

Feature-level Interaction Explanations in Multimodal Transformers

Multimodal Transformers often produce predictions without clarifying how different modalities jointly support a decision. Most existing multimodal explainable AI (MXAI) methods extend unimodal saliency to multimodal backbones, highlighting important tokens or patches within each modality, but they rarely pinpoint which cross-modal feature pairs provide complementary evidence (synergy) or serve as reliable backups (redundancy). We present Feature-level I2MoE (FL-I2MoE), a structured Mixture-of-Experts layer that operates directly on token/patch sequences from frozen pretrained encoders and explicitly separates unique, synergistic, and redundant evidence at the feature level. We further develop an expert-wise explanation pipeline that combines attribution with top-K% masking to assess faithfulness, and we introduce Monte Carlo interaction probes to quantify pairwise behavior: the Shapley Interaction Index (SII) to score synergistic pairs and a redundancy-gap score to capture substitutable (redundant) pairs. Across three benchmarks (MMIMDb, ENRICO, and MMHS150K), FL-I2MoE yields more interactionspecific and concentrated importance patterns than a dense Transformer with the same encoders. Finally, pair-level masking shows that removing pairs ranked by SII or redundancy-gap degrades performance more than masking randomly chosen pairs under the same budget, supporting that the identified interactions are causally relevant. Code is available at https://github.com/dut0817/FL-I2MoE.
Jul 29, 2026cs.CV

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regions or text tokens, they cannot answer a fundamental question: which modality drives a prediction? Consequently, a model may produce the correct output while relying on the wrong source of evidence, masking shortcut learning and unsafe reasoning. We formulate modality attribution as a complementary explainability objective for multimodal foundation models and propose Counterfactual Modality Attribution (CMA), the first framework for quantifying modality-level contributions in MLLMs. CMA generates image-only, text-only, and joint multimodal counterfactuals using coupled diffusion priors and converts them into principled modality attribution scores through a cooperative game-theoretic formulation based on Shapley values. We evaluate CMA on controlled synthetic benchmarks with known ground-truth modality reliance and on a real-world multimodal clinical dataset. CMA correctly identifies the decision-driving modality in 98% of controlled cases and consistently outperforms baselines, revealing failures of cross-modal reasoning that remain invisible to predictive accuracy alone. Our results establish modality attribution as a complementary dimension of explainability beyond feature attribution, providing a principled framework for auditing multimodal foundation models in safety-critical applications.