cs.AIOct 1, 2026

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

Authors: Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia

Organizations: Department of Technologies and Information Systems, Universidad de Castilla-La Mancha, Avenida Carlos III, s/n, Toledo, 45071, Spain · University of Murcia, Avenida Teniente Flomesta, 5, Murcia, 3003, Spain · Murcian Bio-Health Institute (IMIB-Arrixaca), Pabellón Docente del Hospital Clínico Universitario Virgen de la Arrixaca, Murcia, 3120, Spain

Abstract

A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.

Figures & tables

Explore similar work

CardsList
  1. Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions

    May 16, 2026Toshinori Yamauchi, Hiroshi Kera, Kazuhiko KawamotoExplainabilityExplainable AI Methods

  2. Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces

    May 12, 2026Shixing Yu, Promit Ghosal, Kyra GanToken-Level UncertaintyLatent Variable

  3. A Multi-Dimensional Evaluation of Explainability in Media Bias Detection

    Jul 22, 2026Ting Chen, Raina Zhang, Benjamin M. Ampel +1ExplainabilityMulti-Dimensional Evaluation