Period ending 2026-09-14
23 new papers
A weekly snapshot of new work published in Chain-Of-Thought Reasoning.
Twelve weeks of publication activity for this topic as it is defined today.
Weekly history
What was published in this field, kept on the site without email delivery.
Period ending 2026-09-14
A weekly snapshot of new work published in Chain-Of-Thought Reasoning.
Period ending 2026-09-07
A weekly snapshot of new work published in Chain-Of-Thought Reasoning.
Inside this field
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
Within Chain-Of-Thought Reasoning
1,527 papers
verification'' signals are not diagnostic: answer matching observes only the outcome, LLM-as-judge provides subjective and non-verifiable critiques, and scalar rewards (e.g., PRMs/RMs) offer little insight into where a multi-step derivation fails.We propose \textbf{SymDiag}, a neuro-symbolic framework that \textbf{reframes reasoning verification as structured failure diagnosis}. SymDiag translates natural-language CoT into symbolic constraints and performs step-level satisfiability/entailment checks to (i) localize failing steps and (ii) produce verifiable diagnostic evidence, including counterexamples, inconsistency witnesses, and missing-premise indicators. A central challenge is that apparent logic violations'' can be caused either by genuine reasoning defects or by neural-to-symbolic translation noise. SymDiag therefore incorporates a Self-Auditor that disentangles TranslationError from ReasoningError via dual symbolic encodings consistency checks, enabling robust diagnosis under partial observability. Across diverse mathematical, logical, scientific, and general reasoning benchmarks, SymDiag improves detection of unfaithful reasoning and provides substantially more effective feedback for multi-round reasoning repair than outcome-only verification and LLM-based judging, offering a principled foundation for trustworthy and scalable reasoning diagnosis.