Faithfulness of Language Model Explanations

Latest papers 37

All topics
CardsList
  1. Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

    May 19, 2026Somnath Banerjee, Pranav Jha, Rima Hazra +1Multilingual Language Model EvaluationFaithfulness of Language Model Explanations

  2. Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

    May 13, 2026Riya Tapwal, Abhishek Kumar, Carsten MapleLLM-as-a-JudgeFaithfulness of Language Model Explanations

  3. RSAT: Structured Attribution Makes Small Language Models Faithful Table Reasoners

    Apr 30, 2026Jugal Gajjar, Kamalasankari SubramaniakuppusamyTable QALLM Grounding

  4. Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models

    Feb 3, 2026Anish Sathyanarayanan, Aditya Nagarsekar, Aarush RathoreCoT FaithfulnessFaithfulness of Language Model Explanations

  5. A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

    Feb 2, 2026Harry Mayne, Justin Singh Kang, Dewi Gould +3Faithfulness of Language Model ExplanationsLLM Interpretability

  6. Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

    Mar 17, 2025Noah Y. Siegel, Nicolas Heess, Maria Perez-Ortiz +1Faithfulness of Language Model ExplanationsLLM Interpretability