LLM Mathematical Reasoning

LLM: Large Language Model

Momentum

19 papers in the last four weeks, up 375% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 194

All topics
CardsList
  1. ZAYA1-8B Technical Report

    May 6, 2026Robert Washbourne, Rishi Iyer, Tomas Figliolia +15Mixture-of-Experts Language ModelsRL for Language Model Reasoning

  2. Automated Formal Proofs of Combinatorial Identities via Wilf-Zeilberger Guidance and LLMs

    May 6, 2026Beibei Xiong, Hangyu Lv, Junqi Liu +5Automated Theorem ProvingLean Theorem Proving

  3. RAG over Thinking Traces Can Improve Reasoning Tasks

    May 5, 2026Negar Arabzadeh, Wenjie Ma, Sewon Min +1Retrieval-Augmented GenerationLLM Mathematical Reasoning

  4. Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts

    May 1, 2026Sheridan Feucht, Tal Haklay, Usha Bhalla +9Numerical Reasoning in Language ModelsLLM Interpretability

  5. Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

    May 1, 2026Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger +4Mathematical Reasoning BenchmarksLLM Evaluation

  6. When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling

    Apr 29, 2026Zhimin Lin, Yixin Ji, Jinpeng Li +5Test-Time ScalingLLM Routing

  7. QED: An Open-Source Multi-Agent System for Generating Mathematical Proofs on Open Problems

    Apr 27, 2026Chenyang An, Qihao Ye, Minghao Pan +1AI Agents for Scientific DiscoveryAutomated Theorem Proving

  8. MathDuels: Evaluating LLMs as Problem Posers and Solvers

    Apr 23, 2026Zhiqiu Xu, Shibo Jin, Shreya Arya +1Mathematical Reasoning BenchmarksLLM Evaluation

  9. Thinking with Reasoning Skills: Fewer Tokens, More Accuracy

    Apr 23, 2026Guangxiang Zhao, Qilong Shi, Xusen Xiao +3Efficient Language Model ReasoningEfficient Language Model Inference

  10. How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning

    Apr 21, 2026Zhiyuan Zhai, Xinkai You, Wenjing Yan +1LLM Inference EfficiencyRL for Language Model Reasoning

  11. OLLM: Options-based Large Language Models

    Apr 21, 2026Shashank Sharma, Janina Hoffmann, Vinay NamboodiriLLM AlignmentRL for Language Model Reasoning

  12. Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization via Operator Trees

    Apr 21, 2026Xiaoyang Liu, Zineng Dong, Yifan Bai +3LLM Self-CorrectionNeuro-Symbolic Reasoning

  13. MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

    Apr 20, 2026Shaden Alshammari, Kevin Wen, Abrar Zainal +5Mathematical Reasoning BenchmarksMultilingual Language Model Evaluation

  14. Learning to Reason with Insight for Informal Theorem Proving

    Apr 17, 2026Yunhe Li, Hao Shi, Bowen Deng +8Automated Theorem ProvingRL for Language Model Reasoning

  15. Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms

    Apr 17, 2026Tanja Baeumel, Josef van Genabith, Simon OstermannLLM InterpretabilityMechanistic Interpretability

  16. Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning

    Apr 17, 2026Yangyi Fang, Jiaye Lin, Xiaoliang Fu +2RL for Language Model ReasoningLLM Mathematical Reasoning

  17. CoTEvol: Self-Evolving Chain-of-Thoughts for Data Synthesis in Mathematical Reasoning

    Apr 16, 2026Zhuo Wang, Zhuo Zhang, Yafu Li +3Evolutionary OptimizationLLM Training

  18. LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches

    Apr 2, 2026Linyang He, Qiyao Yu, Hanze Dong +5Mathematical Reasoning BenchmarksLLM Evaluation

  19. LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

    Feb 27, 2026Antoine Peyronnet, Fabian Gloeckle, Amaury HayatMathematical Reasoning BenchmarksAutomated Theorem Proving

  20. MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation

    Jan 29, 2026Tianyi Xu, Kosei Uemura, Alfred Malengo Kondoro +6Numerical Reasoning in Language ModelsMathematical Reasoning Benchmarks

  21. The Effect of Scripts and Formats on LLM Numeracy

    Jan 21, 2026Varshini Reddy, Craig W. Schmidt, Seth Ebner +3Numerical Reasoning in Language ModelsMultilingual Language Model Evaluation

  22. MINIF2F-DAFNY: LLM-Guided Mathematical Theorem Proving via Auto-Active Verification

    Dec 11, 2025Mantas Baksys, Stefan Zetzsche, Olivier Bouissou +1Automated Theorem ProvingFormal Verification

  23. ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth

    Dec 2, 2025Salman Rahman, Sruthi Gorantla, Arpit Gupta +3Process Reward ModelsReinforcement Learning with Verifiable Rewards