cs.PLSep 29, 2026

DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis

Authors: Yuan Li, Hanyun Jiang, Guowei Tian, Chengpeng Wang, Peisen Yao

Organizations: The State Key Laboratory of Blockchain and Data Security, Zhejiang University · National University of Singapore

Abstract

Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated. We present DatalogBench, a benchmark of 136 text-to-Datalog synthesis tasks curated from existing Datalog-based artifacts. Synthesized programs are graded by execution on held-out inputs against an oracle validated by mutation analysis. Across six LLMs and four prompting configurations, exact match peaks at 68.4%, and relation descriptions or an input-output example have only modest, model-dependent effects. Under direct prompting, most failures occur at compile time, typically because a model invents auxiliary predicates that it never declares or types consistently. Two coding agents reach up to 83.8% and eliminate nearly all such failures, leaving mostly semantic errors concentrated in recursive tasks. DatalogBench thus identifies recursive reasoning and decomposition as open challenges for current LLMs and agents, and offers a reliable, execution-grounded measure of both.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis

    May 6, 2026Atharva Naik, Yash Mathur, Prakam +2Llm-Driven Code SynthesisSynthetic Task

  2. PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics

    May 29, 2025Atharva Naik, Prakam, Yash Mathur +6Reasoning BenchmarkLarge Language Model Benchmarks

  3. HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs

    Jun 22, 2026Yucheng Wu, Jundong Xu, Mingzhen Ju +4Reasoning Skills