cs.CLAug 10, 2026

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

Authors: Vincent Cohen-AddadDimitris PaparasErnest van WijlandMax SpringerJulien Canitrot-ParadisHonghao LinDavid WoodruffAdarsh Kumarappan+8 more

Organizations: †Google 1 · ‡CNRS, IRIF, Université Paris-Cité · §Princeton University, Department of Computer Science · ¶Université Paris-Saclay, CEA, List, Palaiseau, France · ‖California Institute of Technology, Department of Computer Science

Abstract

We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical computer science venues (STOC, FOCS, and SODA). Each task provides the necessary context to derive a self-contained proof for a target result. We evaluate state-of-the-art models on this benchmark. We verify the correctness of generated proofs via a verification agent, and further benchmark the verifier against human-expert proof judgements on a set of target statements and generated proofs pairs. Our reference verifier achieves over 90% accuracy on the expert labeled set.

Explore similar work

CardsList
  1. Benchmarking Testing in Automated Theorem Proving

    Apr 26, 2026Jongyoon Kim, Hojae Han, Seung-won HwangTheoremProof