cs.AISep 29, 2026

ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning

Authors: İbrahim Ethem Deveci, Funda Tan Çalık, Barış Deniz Sağlam, Duygu Ataman

Organizations: Graduate School of Informatics Middle East Technical University Ankara, Türkiye

Abstract

Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains. Real-world reasoning often proceeds under incomplete and revisable information: conclusions may be supported provisionally, defeated by counter-evidence, reinstated by further arguments, or revised when stronger reasons become available. Reasoning of this kind is generally referred to as defeasible reasoning. We introduce ArgGYM, a procedural benchmark and RLVR-compatible training environment for structured defeasible reasoning. ArgGYM decomposes this reasoning into twelve tasks and grounds task-specific scoring in a symbolic argumentation engine that computes the formal states used to evaluate model outputs. It includes a frozen benchmark of 1,440 verified instances across fifteen curriculum configurations, two argument preference orderings (weakest-link and last-link), and two set orderings (elitist and democratic), while the same generators and verifiers can produce fresh instances for evaluation that reduces dependence on static test sets and for verifiable-reward training. On the frozen benchmark, frontier and open-weight models show sharply different reasoning profiles: they can recover substantial parts of structured answers without solving the complete task, and performance declines in later curriculum configurations with longer dependencies and more interacting structures. We release the benchmark, generators, and verifiers for reproducible evaluation and RLVR training.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

    Aug 31, 2026Jayanta Sadhu, Sayem Shahad, Kenneth MarinoConversational Context

  2. LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening

    May 19, 2026Ming Zhang, Qiyuan Peng, Yinxi Wei +13Reasoning BenchmarkReasoning Skills

  3. DeFAb: A Verifiable Benchmark for Defeasible Abduction in Foundation Models

    Jun 17, 2026Patrick Cooper, Alvaro VelasquezAbductive ReasoningFoundation Model