cs.CLMay 19, 2026

Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

Authors: Somnath BanerjeePranav JhaRima HazraAnimesh Mukherjee

Organizations: Indian Institute of Technology Kharagpur · TCG Crest · National University of Singapore

Abstract

LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model identifies input token spans as evidence alongside a generated rationale'' and uncover a systematic trade-off: English-pivot explanations can achieve higher span agreement with human rationales while their evidence becomes less causally grounded in the model's prediction, as measured by both comprehensiveness and sufficiency. Across 3 tasks, 5languages, and 2multilingual LLM families, we find that English explanations frequently produce fluent but loosely anchored rationales, with comprehensiveness degrading by up to 5.7x relative to native-language conditions - even as task accuracy remains stable across settings. For socially nuanced classification, English pivots also fail to preserve pragmatic cues, reducing both faithfulness and span agreement. We recommend auditing explanations in the input language, reporting multi-faceted faithfulness metrics beyond lexical overlap, and treating English rationales as communication summaries rather than faithful decision traces.

Explore similar work

CardsList
  1. An Empirical Study of Counterfactual Self-Explanations in LLMs

    Sep 15, 2026Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis +2RationalesCounterfactual Explanation

  2. DEPART: DEcomposing PARiTy across Multilingual LLMs

    May 27, 2026Manan Uppadhyay, Prashant Kodali, Pranjal Chitale +3Multilingual Large Language ModelsEnglish