T-MoXAI: A Hierarchical Explainability Framework for Temporal Multimodal Data
Authors: Ali Inha, Mo Vali, Saaliha Vali, Pietro Liò, Meen-Yau Thum
Organizations: Department of Computer Science, University of Cambridge, Cambridge, UK · Cavendish Laboratory, Department of Physics, University of Cambridge, Cambridge, UK · Imperial College Healthcare NHS Trust, London, UK · Lister Fertility Clinic, The Lister Hospital, HCA UK, UK
Artificial Intelligence (AI) models for temporal multimodal data have potential in healthcare and agriculture, but their opacity can limit trust and adoption. We introduce T-MoXAI (Temporal Multimodal eXplainable AI), a hierarchical framework explaining (1) when timepoints influence predictions, using temporal Shapley values; (2) which modalities contribute at those moments, using attention analysis; and (3) what features or image regions drive decisions, using gradient based attribution. A transformer based architecture handles irregular temporal sequences and heterogeneous data, generating all three explanation levels in under one second for interactive decision support. We evaluate the framework on two real world tasks: predicting IVF treatment outcomes from ultrasound sequences and clinical measurements (AUC 0.660 despite significant class imbalance), and forecasting wheat yield from temporal RGB imagery and phenotypic traits (R2 0.265 amid substantial environmental variability). Ablation studies indicate that temporal modelling is critical in both domains: removing it reduces performance to the equivalent of random guessing. Temporal ROAR experiments provide evidence that the explanations reflect the model's reasoning process. With a unified, domain agnostic architecture and open source implementation, T-MoXAI provides a baseline for temporal multimodal XAI, addressing fragmentation in the field and supporting applications where understanding decisions is as important as predictive accuracy.
Figures & tables
Figure 3.1: The T-MoXAI architecture. Raw image and tabular data from each timestep are processed by modality specific encoders (ResNet 50 and MLP). The resulting embeddings are interleaved, combined with positional encodings, and fed into a Transformer Encoder. The final prediction is derived from the [CLS] token’s output representation.