A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
Authors: Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia
Organizations: Department of Technologies and Information Systems, Universidad de Castilla-La Mancha, Avenida Carlos III, s/n, Toledo, 45071, Spain · University of Murcia, Avenida Teniente Flomesta, 5, Murcia, 3003, Spain · Murcian Bio-Health Institute (IMIB-Arrixaca), Pabellón Docente del Hospital Clínico Universitario Virgen de la Arrixaca, Murcia, 3120, Spain
A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.
Figures & tables
Framework / toolkit
Primary purpose
Method families
DeBERTa-v3
Eval metrics
Quantus ( Hedström and others, 2023 )
Evaluation of explanations
Agnostic backends
Partial
Yes
BEExAI ( Sithakoul et al., 2024 )
Benchmark of XAI quality
Agnostic backends
Partial
Yes
Captum ( Kokhlikyan and others, 2020 )
PyTorch attribution API
Gradient, occlusion, LIME, SHAP
Partial
No
Inseq ( Sarti et al., 2023 )
Sequence-model interpretability
Gradient, attention, perturbation
Partial
Basic
Ecco ( Alammar, 2021 )
Visualization of transformers
Attention, gradient
No
No
BertViz ( Vig, 2019 )
Attention visualization
Attention
No
No
Table 1: Comparison of XAI frameworks and toolkits relevant to transformer attribution. Methods : coverage of attribution families. DeBERTa-v3 : native support for disentangled attention. Eval metrics : built-in quantitative evaluation.
Figure 1: Methodology overview. High-level workflow of the five sequential operational stages: corpus balancing, zero-shot NLI inference, XAI methods, Top-20 Jaccard consistency analysis, and SHAP-based error analysis.
Category
Train
Test
Total
Neoplasms
2,530
633
3,163
Digestive
1,195
299
1,494
Nervous
1,540
385
1,925
Cardiovascular
2,441
610
3,051
General
3,844
961
4,805
Total
11,550
2,888
14,438
Table 2: Original distribution of the Medical Abstracts corpus by clinical category.
Category
Selected hypotheses
Neoplasms
Description of a neoplastic condition such as tumor, cancer, malignancy, carcinoma, sarcoma, lymphoma, leukemic process, or metastatic disease; abnormal cell growth consistent with an oncological process; diagnosis, progression, treatment, staging, or complications of a neoplastic disease; oncological process affecting tissues or organs, including primary or metastatic tumors; pathology related to cancer or tumor biology.
Digestive
Disorder of the digestive system affecting the gastrointestinal tract, liver, pancreas, or biliary system; gastrointestinal symptoms, inflammation, infection, obstruction, hepatobiliary disease, or pancreatic pathology; diseases such as gastritis, colitis, pancreatitis, hepatitis, cirrhosis, intestinal dysfunction, or gastrointestinal bleeding; gastroenterological pathology involving digestive organs; digestive disease affecting esophagus, stomach, intestines, colon, liver, or pancreas.
Nervous
Neurological disorder affecting the brain, spinal cord, or peripheral nerves; stroke, neuropathy, seizures, demyelination, meningitis, encephalopathy, or neurodegenerative diseases; neurological deficits, altered consciousness, seizures, focal findings, or neuroinflammatory pathology; pathology affecting neural tissue or circuits; diseases of the central or peripheral nervous system.
Cardiovascular
Cardiovascular disorder affecting the heart or blood vessels; myocardial infarction, heart failure, arrhythmia, hypertension, ischemia, or vascular pathology; cardiological symptoms, hemodynamic changes, perfusion abnormalities, or vascular obstruction; diseases of the heart muscle, coronary arteries, valves, or systemic circulation; cardiovascular pathology affecting cardiac or vascular structures.
General
Systemic or generalized medical condition affecting multiple organs; generalized pathological state with symptoms such as fever, malaise, fatigue, weakness, or systemic inflammation; disorder affecting the body globally rather than a single organ system; diffuse or nonspecific organ anomaly involving general physiological dysfunction; systemic medical problem without a clearly dominant organ system.
Table 3: Enriched hypotheses per clinical category used by the natural language inference engine.
Figure 2: Confusion matrices for DeBERTa-v3 over the balanced sample.Left: counts. Right: normalized percentages. NEO: neoplasms; DIG: digestive; NER: nervous; CAR: cardiovascular; GEN: general.
Figure 3: Architecture for the comparison of explanation methods. Outputs of the five methods are standardized through Top-20 token extraction and compared with the Jaccard index.
Category
Summary of the abstract
Neoplasms
Aggressive facial melanoma with neural invasion, successfully treated by surgical resection, radiotherapy, and reconstructive muscle flap.
Digestive
Optimized hepatic surgical technique using light compression bands to minimize tissue damage and hypoxia during vascular occlusion.
Nervous
Plasticity of the central nervous system in children (3–6 years) to coordinate oropharyngeal musculature after corrective speech surgery.
Cardiovascular
Long-term safety study of verapamil, highlighting its efficacy in blood pressure control and its positive effect on HDL cholesterol without altering the metabolic profile.
General pathological
Reinterpretation of mandibular inflammation as a muscular dysfunction caused by mechanical overload rather than infection, validated by improvement with relaxation techniques.
Table 4: Correctly classified cases used in the explainability comparison. Each abstract is summarized by its dominant clinical content; the full texts are available in the released experimental configuration.
Comparison
Jaccard
SHAP vs. LIME
0.333
SHAP vs. IxG
0.300
SHAP vs. Occlusion
0.500
SHAP vs. AxG
0.333
LIME vs. IxG
0.185
LIME vs. Occlusion
0.200
Table 5: Pairwise Jaccard values for a correctly classified neoplasm case.
Comparison
Jaccard
SHAP vs. LIME
0.352
SHAP vs. IxG
0.200
SHAP vs. Occlusion
0.111
SHAP vs. AxG
0.157
LIME vs. IxG
0.318
LIME vs. Occlusion
0.315
Table 6: Pairwise Jaccard values for a correctly classified digestive case.
Comparison
Jaccard
SHAP vs. LIME
0.200
SHAP vs. IxG
0.368
SHAP vs. Occlusion
0.352
SHAP vs. AxG
0.111
LIME vs. IxG
0.333
LIME vs. Occlusion
0.380
Table 7: Pairwise Jaccard values for a correctly classified nervous system case.
Comparison
Jaccard
SHAP vs. LIME
0.529
SHAP vs. IxG
0.217
SHAP vs. Occlusion
0.250
SHAP vs. AxG
0.150
LIME vs. IxG
0.307
LIME vs. Occlusion
0.238
Table 8: Pairwise Jaccard values for a correctly classified cardiovascular case.
Comparison
Jaccard
SHAP vs. LIME
0.136
SHAP vs. IxG
0.227
SHAP vs. Occlusion
0.059
SHAP vs. AxG
0.188
LIME vs. IxG
0.267
LIME vs. Occlusion
0.036
Table 9: Pairwise Jaccard values for a general pathological case. No token is shared by all five methods.
Comparison
NEO
DIG
NER
CAR
GEN
Mean ± SD
SHAP vs. LIME
0.333
0.352
0.200
0.529
0.136
0.31 ± 0.14
SHAP vs. IxG
0.300
0.200
0.368
0.217
0.227
0.26 ± 0.06
SHAP vs. Occlusion
0.500
0.111
0.352
0.250
0.059
0.25 ± 0.16
SHAP vs. AxG
0.333
0.157
0.111
0.150
0.188
0.19 ± 0.08
LIME vs. IxG
0.185
0.318
0.333
0.307
0.267
0.28 ± 0.05
LIME vs. Occlusion
0.200
0.315
0.380
0.238
0.036
0.23 ± 0.12
Table 10: Pairwise Jaccard index per category and method pair. NEO: neoplasms; DIG: digestive; NER: nervous; CAR: cardiovascular; GEN: general.
Figure 4: Architecture for the error analysis protocol of the general pathological category. SHAP attributions are visualized through saliency maps and waterfall plots.
Case
Predicted
Top driving tokens
Failure pattern
1
Cardiovascular
hemodynamic , atherosclerosis , glomerular
SO
2
Neoplasms
malignancy , tumor , carcinoma
LH
3
Cardiovascular
pressure , blood
SO
4
Neoplasms
it , diagnose (no robust descriptor)
LAC
5
Neoplasms
biopsy
LH
Table 11: Summary of the five misclassified general pathological cases inspected with SHAP. Predicted : erroneous category assigned by DeBERTa-v3. Top driving tokens : tokens with the largest positive SHAP attribution per case. Failure pattern : LH lexical hypersensitivity, SO semantic overlap, LAC loss of attribution coherence (see text).
Case
Predicted
Summary of the abstract
1
Cardiovascular
Pathophysiology of glomerular damage and scarring in chronic kidney disease, highlighting the interaction between genetic and metabolic factors.
2
Neoplasms
Experimental study on the malignant transformation of keratinocytes infected with HPV-18, linking tumour progression to chromosomal instability and cell-cycle failures.
3
Cardiovascular
Investigates platelet dysfunction in uraemic patients, identifying excess nitric oxide as the key mediator of bleeding and its therapeutic reversibility.
4
Neoplasms
Studies streptococcal pharyngitis, its treatment with penicillin, and the systemic risks derived from inadequate management of the bacterial infection.
5
Neoplasms
Safety evaluation of outpatient lung biopsy in 169 patients, showing a low incidence of severe complications and the feasibility of same-day discharge.
Table 12: Clinical content of the five misclassified general pathological cases. Predicted : erroneous category assigned by DeBERTa-v3. The true category in all cases is General pathological conditions .
Figure 5: Error analysis of a general case misclassified as cardiovascular. Left: saliency map. Right: SHAP waterfall plot.
Figure 6: Error analysis of a general case misclassified as neoplasm. Left: saliency map. Right: SHAP waterfall plot.
Zero-shot textual explanations aim to make image classifiers more transparent by probing their internal representations, without relying on task-specific supervision or LVLMs. However, existing methods often miss the features that truly drive the prediction, resulting in limited \textit{faithfulness} to the evidence underlying the model's decision. To address this, we propose FaithTrace. Motivated by the idea that faithful explanations should describe concepts that strongly influence the prediction, FaithTrace directly measures how much the representation induced by the explanation changes the class logit. We introduce an influence score, computed as the directional derivative of the class logit along the text-induced direction in the classifier's feature space, and use it as a proxy for faithfulness. Moreover, we extend this influence score into quantitative evaluation metrics, helping fill the gap in faithfulness evaluation for textual explanations. Experiments show that FaithTrace yields more faithful explanations than baselines, facilitating a more accurate understanding of the model. The code will be publicly released.
Toshinori Yamauchi, Hiroshi Kera, Kazuhiko Kawamoto
Chiba University · National Institute of Informatics
A critical step for reliable large language models (LLMs) use in healthcare is to attribute predictions to their training data, akin to a medical case study. This requires token-level precision: pinpointing not just which training examples influence a decision, but which tokens within them are responsible. While influence functions offer a principled framework for this, prior work is restricted to autoregressive settings and relies on an implicit assumption of token independence, rendering their identified influences unreliable. We introduce a flexible framework that infers token-level influence through a latent mediation approach for general prediction tasks. Our method attaches sparse autoencoders to any layer of a pretrained LLM to learn a basis of approximately independent latent features. Unlike prior methods where influence decomposes additively across tokens, influence computed over latent features is inherently non-decomposable. To address this, we introduce a novel method using Jacobian-vector products. Token-level influence is obtained by propagating latent attributions back to the input space via token activation patterns. We scale our approach using efficient inverse-Hessian approximations. Experiments on medical benchmarks show our approach identifies sparse, interpretable sets of tokens that jointly influence predictions. Our framework enhances trust and enables model auditing, generalizing to high-stakes domain requiring transparent and accountable decisions.
Shixing Yu, Promit Ghosal, Kyra Gan
Electrical and Computer Engineering · Cornell Tech · Department of Statistics +2
Detecting media bias automatically is difficult because biased framing is often subtle, yet in domains such as news analysis, accurate predictions alone are insufficient without explanations that reflect the model's underlying reasoning. We present a multi-dimensional evaluation of explainability in encoder-based media bias detection using the Bias Annotations By Experts (BABE) dataset. Specifically, we study BERT and RoBERTa as classifiers (base and large variants) along three complementary axes: predictive performance, explanation plausibility (token-level alignment with expert rationales), and mechanistic faithfulness (whether compact sets of attention heads recover predictive signal under counterfactual rationale masking). To induce variation in plausibility, we additionally investigate attention-supervised finetuning, which incorporates expert rationale annotations as an auxiliary training signal. Attention supervision serves as an intervention on attribution plausibility, while the effectiveness of attribution methods varies substantially across architectures. Circuit analysis further reveals substantial variation in mechanistic recoverability across architectures, suggesting that model scale alone does not determine circuit compressibility. Taken together, our findings suggest that predictive performance, attribution plausibility, and mechanistic faithfulness characterize different aspects of model behavior and should be evaluated separately when studying explainability in media bias detection.
Ting Chen, Raina Zhang, Benjamin M. Ampel +1
Carnegie Mellon University · Indiana University · Georgia State University