Theorem Proving

Momentum

24 papers in the last four weeks, up 243% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 172

All topics
CardsList
  1. AIProver: Agentic Auto-Formalization of Mathematical Research via Certificate-Driven Evolving Harness

    Oct 4, 2026Prithwish Jana, Viet Bach Hoang, Logan Luna +9Theorem Proving

  2. LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification

    Oct 1, 2026Haochen Zhang, Laura Yao, Zachary Plotkin +2Video CaptioningReinforcement Learning With Verifiable Reward

  3. Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

    Oct 1, 2026Zihan Liu, Xurong XieMathematical ReasoningReinforcement Learning With Verifiable Reward

  4. Cogentic: Multi-Agent Orchestration for Automated Proof Discovery

    Sep 30, 2026Yang Cai, Vineet Gupta, Yanchen Jiang +4Theorem ProvingMulti-Agent Orchestration

  5. Growing an Agent/Prover Interface: Evolutionary Tool Design for Cost-Efficient Theorem Proving in Rocq and Lean

    Sep 30, 2026Jules Viennot, Guillaume Baudart, Marc LelargeTheorem ProvingMath Problems

  6. LeanPolish: Verified Supervision for Lean Proof Compression

    Sep 29, 2026Pauline BourigaultTheorem ProvingProof

  7. Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

    Sep 28, 2026Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang +7Credit AssignmentAgentic Reinforcement Learning

  8. TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science

    Sep 28, 2026Chutong Yang, Xiyuan Zhang, Yu Huang +7Theorem ProvingProof

  9. When Does Structured Knowledge Help Neural Theorem Proving?

    Sep 28, 2026Sareh Nabi, Roland Vogl, Marzieh NabiTheorem Proving

  10. Learning to Discover Interesting Mathematics

    Sep 23, 2026Niket Patel, Ahmad Rammal, Amaury Hayat +2Theorem Proving

  11. Direct Optimization of Generators for Search in Automated Theorem Proving

    Sep 22, 2026Adam Ousherovitch, Ambuj TewariTheorem ProvingCross-Entropy Losses

  12. SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

    Sep 18, 2026George Ma, Benjamin Mikek, Haoyu Li +9Swe-Bench VerifiedFormal Verification

  13. Structured Four-Stage Legal Translation: From Natural-Language Traffic Rules to PROLOG

    Sep 17, 2026May Myo Zin, Wachara Fungwacharakorn, Ken Satoh +1Legal Reasoning TasksNatural Language

  14. Sage: Formalization with Semantic Correction

    Sep 16, 2026Thomas Hirtz, Farzad Jafarrahmani, Abdelmouksit Sagueni +3AutoformalizationTheorem Proving

  15. Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

    Sep 14, 2026Honghao Lin, David P. Woodruff, Yuan Deng +3Theorem ProvingMathematics

  16. Gödel's and Scott's Variants of the Ontological Argument in Lean 4 and TPTP THF

    Sep 14, 2026Christoph BenzmüllerTheorem ProvingProof

  17. A machine-checked proof of the Dong-Yang classification of optimal (n,4) binary codes for BSCs

    Sep 12, 2026Shenghao Yang, Yanyan DongSource-Channel CodingBits

  18. Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification

    Sep 12, 2026Joshua Ong Jun Leang, Haonan Li, Zheng Zhao +6Mathematical ReasoningTheorem Proving

  19. StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

    Sep 8, 2026Idan Davidovich, Debargha Ganguly, Vikash Singh +1Theorem ProvingStochastic Processes

  20. Monadic Second-Order Logic in HOL: Deep and Shallow with Automated Faithfulness (Extended Preprint)

    Sep 7, 2026Christoph Benzmueller, Daniel KirchnerFirst-Order LogicTheorem Proving

  21. AutoGraphForge: Towards Automated Graph Theory Discovery

    Sep 3, 2026Ján PastorekTheorem ProvingOpen Problems

  22. A Certificate-Producing Cascade for Equational Implication: The SAIR EQT2 Stage 2 Solver

    Sep 1, 2026Haobo Ma, Wenlin Zhang, Manuel Israel CázaresEquivalenceTheorem Proving

  23. Prove2Me: An Open Collaborative Platform for Scaling Math Formalization

    Aug 28, 2026Shuze Chen, Kunal Marwaha, Xiaoyang Lu +2Theorem ProvingAutoformalization

  24. CAPRI: Contract-Aware Proof Repair for Isabelle

    Aug 13, 2026Jim Woodcock, Gabriel Leite, Augusto Sampaio +1Theorem ProvingProof

  25. PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs

    Aug 13, 2026Sadat Shahriyar, Shareef Ahmed, Abdullah Al ArafatTheorem ProvingSchedulers

  26. FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation

    Aug 11, 2026Rob Cornish, Iacopo Ghinassi, Po-Hung Yeh +7AutoformalizationExplanation Faithfulness

  27. NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering

    Aug 6, 2026Jonas Gann, Michael GertzHievi-RagSearch-Augmented Reasoning

  28. Can Open-Weight LLMs Produce Kernel-Verified Coq Proofs? A Pilot Study

    Aug 5, 2026Ahmed Ryan, Md Erfan, Akond Ashfaque Ur Rahman +1Theorem ProvingProof

  29. PPDL: LLM-Based Flows as Probabilistic Programs

    Aug 5, 2026Louis Mandel, Guillaume Baudart, Mandana Vaziri +1Large Language Model UncertaintyLarge Language Model Workflows

  30. Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates

    Jul 31, 2026Bohan Chen, Shivam N. Patel, Richard Hoffmann +2Symbol-Theorem Proving

  31. LeanCSP: A Framework for Certifying Constraint Reformulation and Solving in Lean

    Jul 30, 2026Pablo Manrique, Stefan SzeiderSatisfiabilityTheorem Proving

  32. BlueprintRepair: Typed Local Edits for Failed Lean Proof Blueprints

    Jul 30, 2026Ruslan KhrulevTheorem ProvingRepair

  33. Distilling Answer Set Programming Theories from Large Language Models

    Jul 30, 2026Nelson Higuera Ruiz, Markus Hofmarcher, Claudiu Leoveanu-CondreiAnswer Set ProgrammingNeuro-Symbolic Framework

  34. MECA: A Mechanism-Centered Agent for Constructing Well-Specified and Valuable Mathematical Conjectures

    Jul 30, 2026Wentao Long, Yunfei Zhang, Chenyi Li +1Open ProblemsTheorem Proving

  35. Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration

    Jul 30, 2026Ting Gong, Michael Ruofan Zeng, Yong YangResearch-Level MathematicsTheorem Proving

  36. Formalizing Flag Algebras in Lean

    Jul 26, 2026Gyeongwon Jeong, Seonghun Park, Jihoon Hyun +2Acyclic GraphsAlgebraic Structures

  37. Learned Interventions in Lean 4 grind

    Jul 25, 2026Evan Wang, Simon Chess, Sophie Szeto +1Theorem ProvingHeuristics

  38. The Best Programming Language for Tokenmaxxing: An Investigation of Coding Agent Behavior Across Programming Languages

    Jul 24, 2026Zixuan Wu, Carolyn Jane Anderson, Arjun GuhaMultilingual AgentsCoding Agents

  39. Case study: solving P-99 with LPTP and an LLM

    Jul 23, 2026Fred Mesnard, Thierry Marianne, Étienne Payet +1Theorem ProvingCase Study

  40. Animation, Verification and Visualisation of Prolog Transition Systems with ProB

    Jul 23, 2026Jan Gruteser, Michael Leuschel, Katharina Engels +1Probabilistic Model CheckingTheorem Proving

  41. Encoding Event-B Proof Rules in Prolog: An Interactive Sequent Prover for ProB

    Jul 23, 2026Katharina Engels, Jan Gruteser, Michael LeuschelTheorem ProvingProof

  42. Case study: proving sqrt(2) irrational with LPTP and an LLM

    Jul 23, 2026Fred Mesnard, Étienne Payet, Wim VanhoofTheorem ProvingFirst-Order Logic

  43. Agree on the Model, Verify the Inference: GKR Protocols for HND-Based Transformer Inference

    Jul 23, 2026Xiaolong Liang, Juanjuan Li, Rui Qin +1Homomorphic EncryptionZero-Knowledge Proofs

  44. AoA: Theorem Proving Agent over Abstract Syntax Tree of Redesigned Language

    Jul 17, 2026Qiyuan Xu, Joshua Ong Jun Leang, Renxi Wang +4Theorem ProvingProof

  45. MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research

    Jul 16, 2026Junjie Zhang, Jiayu Liu, Wenbin Liu +11Theorem ProvingHuman-Ai Collaboration

  46. AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

    Jul 13, 2026Lingkai Kong, Zijian Wu, Yuzhe Gu +11Research-Level MathematicsMathematical Reasoning Benchmarks

  47. Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization

    Jul 13, 2026Tian-Shuo Liu, Shiyuan Zhang, Zijie Geng +5AutoformalizationTheorem Proving

  48. TreeThink: A Modular Tree Search Library for Mathematical Reasoning with LLMs

    Jul 13, 2026Burak S. Akbudak, Zeynel A. Uluşan, Can S. Erer +1Theorem ProvingTree Search

  49. First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint)

    Jul 12, 2026Christoph Benzmüller, Daniel KirchnerFirst-Order LogicLogic

  50. OpenProver: Agentic and Interactive Theorem Proving with Lean 4

    Jul 10, 2026Matěj Kripner, Milan StrakaTheorem ProvingFormal Verification

  51. Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

    Jul 7, 2026Pavel Snopov, German MagaiTheorem ProvingLarge Language Model Agents