cs.CLOct 8, 2026

Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog

Authors: Hadeel Al-Negheimish, Jasna Ilieva, Yoon Kim

Organizations: King Saud University · Massachusetts Institute of Technology

Abstract

Current frontier LLMs can theoretically process long contexts with 1M tokens or more. But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts? We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog. We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length. We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key

    May 7, 2026Tianle Wang, Zhaoyang Wang, Guangchen Lan +4Logical ReasoningRL for Language Model Reasoning

  2. PrologMCP: A Standardized Prolog Tool Interface for LLM Agents

    Jun 12, 2026Agnieszka Mensfelt, Adarsh Prabhakaran, Adrian Haret +2Logic ProgrammingSymbolic Reasoning

  3. Recursive Models for Long-Horizon Reasoning

    Mar 2, 2026Chenxiao Yang, Nathan Srebro, Zhiyuan LiLong-Context ModelingLanguage Modeling