cs.AISep 29, 2026

XU-RS: Explaining Credal Width in Random-Set Language Models

Authors: David Achara, Maryam Sultana, Alexander D. Rast, Fabio Cuzzolin

Organizations: Institute for AI, Data Analysis and Systems (AIDAS) Oxford Brookes University

Abstract

Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained language models. These classifiers assign probability to individual answers and to groups of answers, producing lower and upper probabilities for each answer; The difference between these probabilities, called credal width, is used to represent epistemic uncertainty about an answer arising from limited training data. We propose XU-RS, a framework that attributes an answer's credal width to the input tokens (words or word pieces) supplied to a language model. XU-RS uses Expected Gradients (a standard feature attribution method) to estimate how input tokens contribute to credal width. The proposed framework is evaluated on a MedQA dataset using SmolLM3-3B and Llama-2-7B models, demonstrating that setting the embedding of a token ranked highly by XU-RS to zero (zero-masking) causes larger changes in credal width than zero-masking randomly selected tokens. In addition, we show that normalisation can cause other answer groups to influence an answer's width, reveal how token attribution can mask numerical errors, and provide diagnostic checks to verify whether a token ranked highly by XU-RS meaningfully explains model uncertainty.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reading Calibrated Uncertainty from Language Model Trajectories

    May 19, 2026Aliai Eusebi, Alexander Herzog, Xiaoyu Liang +3Large Language Model UncertaintyCalibrated Uncertainty

  2. Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

    Sep 29, 2026Feiyang Li, Shengjing Liu, Qi Zhan +7Large Language Model UncertaintyConfidence Calibration

  3. Credal Large Language Models for Semantic Commitment under Uncertainty

    Aug 24, 2026Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio CuzzolinLarge Language Model UncertaintyCryptographic Merkle Tree Commitments