Faithfulness of Language Model Explanations

Latest papers 37

All topics
CardsList
  1. A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

    Oct 1, 2026Zhangshu Joshua Jiang, Zina Ibrahim, James T. TeoFaithfulness of Language Model ExplanationsTemporal Reasoning in Language Models

  2. How the Audit Rule Shapes Faithful Factor Explanations in LLMs

    Oct 1, 2026Taolin Zhang, Hanyu Wang, Jiuheng Wan +2Counterfactual ExplanationsFaithfulness of Language Model Explanations

  3. When Do Biological Reasoning Models Use Their Biological Inputs?

    Oct 1, 2026Ada Fang, Nikitha Thoduguli, Lukas Fesser +3Faithfulness of Language Model ExplanationsCausal Interventions in Language Models

  4. Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions

    Sep 30, 2026Sujung Kim, Seung Hwan Cho, Sangjin Park +1Numerical Reasoning in Language ModelsFaithfulness of Language Model Explanations

  5. Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces

    Sep 29, 2026Ratish Puduppully, Pranabendu Misra, Paarth Iyer +3Faithfulness of Language Model ExplanationsCoT Evaluation

  6. Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

    Sep 28, 2026Li Zhang, Chuqin Geng, Mark Zhang +4Faithfulness of Language Model ExplanationsMechanistic Interpretability

  7. When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment

    Sep 28, 2026Zhaohan Zhang, Junjie Liu, Chengzhengxu Li +6Faithfulness of Language Model ExplanationsConfidence Estimation in Language Models

  8. Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

    Sep 27, 2026Gert Lek, Zixuan Xia, Pin-Yu Chen +1Transformer InterpretabilityFaithfulness of Language Model Explanations

  9. EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing

    Sep 16, 2026Sihao Ding, Santosh Vasa, Aditi Ramadwar +1VLM EvaluationVLM Robustness

  10. An Empirical Study of Counterfactual Self-Explanations in LLMs

    Sep 15, 2026Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis +2Counterfactual ExplanationsFaithfulness of Language Model Explanations

  11. Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an Oracle

    Aug 31, 2026Dennis Gross, Helge SpiekerLLM EvaluationFaithfulness of Language Model Explanations

  12. Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk

    Aug 8, 2026Gregorius Reynaldi Pratama, Kuo-Kun TsengFeature AttributionFaithfulness of Language Model Explanations

  13. Training Large Language Models for Self-Explanation Faithfulness

    Jul 23, 2026Yeoktatt Cheah, María Pérez-Ortiz, Noah Y. Siegel +1Faithfulness of Language Model ExplanationsLLM Interpretability

  14. CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness

    Jul 21, 2026Ziming Wang, Yinghua Yao, Changwu Huang +2CoT FaithfulnessFaithfulness of Language Model Explanations

  15. From Plausible to Actionable: A Position on LLM Self-Explanations

    Jul 17, 2026Elize Herrewijnen, Benedetta Muscato, Gizem Gezici +1CoT FaithfulnessFaithfulness of Language Model Explanations

  16. Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers

    Jun 19, 2026Vatsal Ananthula, Adarsh KumarappanFaithfulness of Language Model ExplanationsLLM Auditing

  17. CAREF: Calibration-Aware Regularization for Explanation Faithfulness Without Rationale Supervision

    May 27, 2026Naphat Nithisopa, Teerapong PanboonyuenFaithfulness of Language Model ExplanationsLLM Interpretability

  18. Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

    May 24, 2026Yoav Gur-Arieh, Ana Marasović, Mor GevaCoT FaithfulnessLLM Evaluation