CoT Faithfulness

CoT: Chain-of-Thought

Momentum

5 papers in the last four weeks, up 67% on the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 54

All topics
CardsList
  1. When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

    Jun 9, 2026Sai Kartheek Reddy Kasu, Nils Lukas, Samuele PoppiCoT FaithfulnessLLM Alignment

  2. The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

    May 27, 2026Eric Onyame, Runtao Zhou, Kowshik Thopalli +2CoT FaithfulnessMultilingual Language Model Evaluation

  3. GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought

    May 26, 2026Weijiang Lv, Wentong Zhao, Jiayu Wang +3CoT FaithfulnessRL for Language Model Reasoning

  4. Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy

    May 25, 2026Xu Shen, Zhen Tan, Song Wang +4CoT FaithfulnessLLM Auditing

  5. Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

    May 24, 2026Yoav Gur-Arieh, Ana Marasović, Mor GevaCoT FaithfulnessLLM Evaluation

  6. Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

    May 24, 2026Jingyi Sun, Qianli Wang, Pepa Atanasova +2CoT FaithfulnessContextual Faithfulness

  7. Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

    May 22, 2026Jinghan Jia, Joe Benton, Eric EasleyCoT FaithfulnessLLM Interpretability

  8. Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models

    May 17, 2026Nicanor Mayumu, Xiaoheng Deng, Patrick MukalaCoT FaithfulnessVLMs for Autonomous Driving

  9. When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel

    May 12, 2026Wenkai Li, Fan Yang, Ananya Hazarika +2CoT FaithfulnessLLM Auditing

  10. Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning

    Apr 23, 2026Qinan Yu, Alexa Tartaglini, Peter Hase +2CoT FaithfulnessReinforcement Learning with Verifiable Rewards

  11. AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency

    Apr 17, 2026Max Henning Höth, Kristian Kersting, Björn Deiseroth +1CoT FaithfulnessFeature Attribution

  12. Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models

    Apr 16, 2026Danae Sánchez Villegas, Samuel Lewis-Lim, Nikolaos Aletras +1CoT FaithfulnessVision-Language Models

  13. MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

    Mar 30, 2026Han Wang, Yifan Sun, Brian Ko +8CoT FaithfulnessLLM Evaluation

  14. ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs

    Mar 19, 2026Abhinaba Basu, Pavan ChakrabortyExplanation EvaluationLLM Interpretability

  15. Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models

    Feb 3, 2026Anish Sathyanarayanan, Aditya Nagarsekar, Aarush RathoreCoT FaithfulnessFaithfulness of Language Model Explanations

  16. Do Language Models Reason Across Languages?

    Jan 10, 2026Yan Meng, Wafaa Mohammed, Christof MonzCoT FaithfulnessCross-Lingual Reasoning in Language Models

  17. Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations

    Apr 7, 2025Pedro Ferreira, Wilker Aziz, Ivan TitovCoT FaithfulnessReward Hacking