cs.CLMay 26, 2026

Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids

Authors: Fabian LukassenJan HerrmannChristoph WeisserAlexander SilbersdorffBenjamin SaefkenThomas Kneib

Organizations: University of Göttingen · BASF SE · Hochschule Bielefeld · TU Clausthal

Abstract

Prior work shows that Large Language Models (LLMs) can transform Explainable AI (XAI) outputs into Natural Language Explanations (NLEs) that score highly on quality metrics such as plausibility, coherence, and comprehensibility. But does explanation quality translate to practical usefulness? We investigate this question in a time-series energy forecasting domain through five controlled experiments (2,730 judgments across 60 test instances), each operationalising a distinct facet of usefulness studied in the XAI literature. Holding NLE quality constant at the high levels established by a prior factorial study, we find that NLEs do not improve task accuracy on any of the five tasks, while inflating self-reported confidence. A placebic control shows that this confidence boost is driven by text presence rather than content. In an out-of-distribution detection task, NLEs reduce the LLM judge's ability to flag unreliable predictions, providing false reassurance that masks model failure. We characterise these findings as the Quality-Usefulness Gap and argue that evaluation of the XAI-to-NLE pipeline must extend beyond text-quality metrics to downstream task performance.

Explore similar work

CardsList
  1. XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

    Sep 10, 2026Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein +1XaiExplainable Artificial Intelligence

  2. Human Decision-Making with Persuasive and Narrative LLM Explanations

    May 22, 2026Laura R. Marusich, Mary Grace Kozuch Dhooghe, Jonathan Z. Bakdash +1Large Language Model Decision-MakingPersuasion