Organizations: Department of Computer Science, Tulane University · Department of Biochemistry and Molecular Biology, Tulane University School of Medicine
Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-epitope prediction models, TCR-SRIM, TULIP, MixTCRpred, and NetTCR-2.2, spanning post-hoc and interpret-by-design approaches as well as transformers and CNNs. To investigate how structural information affects model explanations, we introduce a benchmark, TCR-XAI2, containing 388 unique experimentally resolved TCR-epitope structures, complemented by structures predicted using AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Using EFT with TCR-XAI2, we demonstrate that (1) CNN and transformer models exhibit distinct learning trajectories; (2) TCR α and β evidence can conflict during learning, limiting the benefits of jointly modeling both chains, while MHC information mitigates this; and (3) real versus predicted structural data for TCR-epitope prediction exhibits distinct TCR and peptide feature preferences as well as differing trajectories of model certainty.
Figures & tables
Figure 1: The EFT paradigm tracks model interpretations during training to obtain importance trajectories, clusters them into interpretation trajectories, and organizes them into learning trajectories that provide explanations. The diagram of TCR-epitope binding is from Parizi et al. (2026) .
Figure 2: EFT explanations reveal distinct learning trajectories across CNN and transformer and feature fusion strategies, while showing that models struggle to learn TCR α and β simultaneously.
Figure 3: A EFT case study of TULIP-NoMHC on a human cytomegalovirus sample reveals conflicts between TCR α and β information for model learning during training.
Figure 4: EFT explanations for TULIP-MHC and MixTCRpred trained on samples with only the HLA-A*02 allele show that MHC information can reduce conflicts between TCR α and β information and enable models to learn features from more chains.
Figure 5: The evolution of BRHR@.25 for CDR3a, CDR3b, and peptide during TULIP training with and without MHC allele information quantitatively demonstrates that MHC information can reduce conflicts between TCR α and β information.
Figure 6: EFT explanations and evolution of BRHR@.25 during training for TCR-SRIM regularized with experimentally resolved structures and AlphaFold3-predicted structures. Structural information rapidly stabilizes model interpretations, while experimentally resolved and predicted structures lead to different interpretation trajectories during training.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: The statistics of sample distribution of TCR-XAI2 benchmark. The TCR-XAI2 benchmark consists of 388 unique structures and 619 samples, while a structure can provide multiple TCR-epitope copies. Most of the samples are from human (80.8%), MHC-I (75.8%), and resolved by X-ray diffraction (98.9%).
Figure 8: The sequence and structural differences between the training and test splits of the TCR-XAI2 benchmark. From the sequence perspective, the median distance between the training and test splits is 18 residues, with a minimum distance of 5 residues. From the structural perspective, the CDR3-peptide contact maps exhibit a median difference of at least 2.65 Åbetween the training and test splits.
Model
Train
Test
Total
Missing
TCR-XAI2 (AlphaFold3)
303
301
604
15
TCR-XAI2 (Boltz-2)
310
309
619
0
TCR-XAI2 (OpenFold)
310
309
619
0
TCR-XAI2 (TCRModel2)
298
293
591
28
TCR-XAI2 (tFold-TCR)
310
309
619
0
TCR-XAI2
310
309
619
-
Appendix
Table 1: The number of samples and training and test splits in the predicted TCR-XAI2 benchmarks. AlphaFold3 and TCRModel2 failed to predict 15 and 28 samples, respectively, while all other methods successfully predicted structures for all samples.
Predictor
CDR3a-peptide (Å)
CDR3b-peptide (Å)
Full (Å)
Boltz-2
2.48±7.78
2.38±7.59
3.24±5.88
AlphaFold3
2.97±7.88
2.95±7.74
3.75±6.01
TCRmodel2
3.38±7.98
3.35±7.74
-
tFold-TCR
6.35±7.54
7.28±6.89
10.29±7.48
OpenFold3
23.47±13.56
16.47±8.14
35.44±3.30
Appendix
Table 2: The structural quality of predictions from different models. We compare the mean RMSD and standard deviation of the full structure and CDR3-peptide interfaces between predicted and experimentally resolved structures. Full-structure RMSD is not available for TCRModel2 because it omits several components of the complete TCR-pMHC complex.
Model
ROC-AUC@FPR <0.1
ROC-AUC
IMMREP23
IMMREP25
Original Test
Original Test
MixTCRpred
0.611
0.513
0.746
0.850
NetTCR-2.2
0.575
0.516
0.748
0.886
TULIP-MHC
0.574
0.745
0.607
0.710
TULIP-NoMHC
0.587
0.744
0.632
0.736
TCR-SRIM
0.580
0.592
0.624
0.641
Appendix
Table 3: The final performance comparison of all models on IMMREP23, IMMREP25, and their respective original test datasets. ROC-AUC at a false positive rate below 0.1 (ROC-AUC@FPR <0.1 ) indicates that all models converged to performance comparable to that reported in their original publications during our retraining.
Figure 9: Different random seeds run for TULIP-NoMHC with QCAI and NetTCR-2.2 with GradCAM.
Figure 10: EFT explanation for MixTCRpred interpreted by GradCAM and AttnLRP respectively.
Figure 11: EFT explanation for MHC-TULIP interpreted by QCAI and AttnLRP respectively.
Figure 12: The EFT explanation for MixTCRpred interpreted by AttnLRP on HLA-A02 , HLA-A11 , and HLA-A*24 specific subset.
Figure 13: The evolution of BRHR@0.25 during training for TCR-SRIM without structural regularization. The model’s interpretation remains highly unstable throughout training, with substantial fluctuations in BRHR across the three interaction regions.
Figure 14: The evolution of BRHR@0.25 during training for TCR-SRIM regularized with experimentally resolved structures and structures predicted by AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Structural regularization stabilizes the model’s interpretation of CDR3-peptide interactions, with the degree of stabilization varying according to the quality of the structural guidance.
Figure 15: EFT for TCR-SRIM regularized with experimentally resolved structures and structures predicted by AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3.
T cell receptor (TCR)-epitope binding prediction is essential for understanding adaptive immunity and developing immunotherapies. Existing sequence- and structure-based models often generalize poorly to unseen epitopes and provide limited interpretability. Furthermore, the impact of generated structures on model learning remains unclear. We present TCR-SRIM, a structure-regularized interpretable-by-design model that combines protein language model embeddings with interpretable contact prototypes to capture residue-level TCR-epitope interactions. TCR-SRIM achieves state-of-the-art predictive performance and improved interpretation quality on the TCR-XAI benchmark. Using its inherent interpretability, we further evaluate the effect of generated structures on model learning. While structures predicted by AlphaFold3, TCRModel2, and tFold-TCR yield competitive performance, they lead to less accurate interaction patterns and reduced binding-site diversity than experimentally-resolved structures. Our results highlight limitations of current structure prediction models for TCR-epitope learning and demonstrate the value of interpretable-by-design models for studying generated biological structures.
Jiarui Li, Zixiang Yin, Yunbei Zhang +4
Department of Computer Science, Tulane University · Department of Biochemistry and Molecular Biology, Tulane University School of Medicine
Accurate computational prediction of T cell receptor (TCR) antigen specificity would transform the study of T cell biology and enable scalable immune engineering, yet existing models lack sufficient sensitivity and specificity for broad applications. A major limitation is the absence of rigorously defined, unseen benchmark datasets that allow unbiased evaluation of model performance and generalizability. Here, we describe two complementary classes of datasets that meet this criterion and argue that they provide both a robust framework for model assessment and a foundation for next-generation TCR-antigen prediction algorithm development.
Yiming Liao, Yiheng Li, Ning Jiang +2
Trustworthy and Intelligent Computing Lab (TAIC), Department of Computer Science and Electrical Engineering, University of Maryland, Baltimore County, Baltimore, Maryland 21250, USA · Children’s Hospital of Philadelphia, Philadelphia, PA 19104, USA · Department of Bioengineering, University of Pennsylvania, Philadelphia, PA 19104, USA +5
Machine learning models are primarily judged by predictive performance, especially in applied genomics, where explanations are read as biological findings. In practice, reported gene panels are stabilised by averaging, ranking, or taking consensus over the many models a pipeline produces across cross-validation folds, tuning grids, and repeated runs. This raises an overlooked question: when two models achieve high accuracy, do they rely on the same internal logic, or reach the same outcome via different mechanisms? We introduce EvoXplain, a diagnostic framework that measures whether a pipeline's explanation is uniquely determined across repeated training and model selection. Rather than analysing a single trained model, EvoXplain treats explanations as samples drawn from the training and model selection pipeline itself, without aggregating predictions or constructing ensembles, and examines whether they form a single coherent explanatory basin or separate into multiple structured basins. We evaluate EvoXplain on a TCGA pan-cancer cohort and a within-cancer breast-cancer subtype task, using elastic-net Logistic Regression and gradient-boosted trees. Although all models reach about 98% accuracy, explanation structure differs across pipelines. Holding the data split fixed and varying only the regularisation strength, equally accurate Logistic Regression models separate into a few discrete, reproducible basins that recur across 100 data splits and carry distinct biological content, while the gradient-boosted pipeline converges to one basin. The same multiplicity appears within a single cancer subtype, from the ordinary tuning step alone. EvoXplain makes explanatory structure visible, revealing when an averaged consensus corresponds to no single trained model, and reframes interpretability as a property of the training pipeline rather than of any single model.