Declarative logic programs offer a powerful and interpretable abstraction for encoding relational structure and neurosymbolic reasoning, by expressing dependencies as weighted compositional rules. However, inducing them from data remains fundamentally hard, bottlenecked by the combinatorial explosion of symbolic search spaces. LLMs have recently emerged as powerful hypothesis generators, but when used in isolation, they lack the capacity to do systematic inductive reasoning needed to reliably synthesize valid programs that fit complex relational distributions. We introduce grasp (Gradient-boosted Synthesis of Probabilistic logic programs), a neurosymbolic framework that casts relational structure learning as functional gradient boosting in which the weak learner is a first-order rule and the intractable inner search is delegated to an LLM proposal oracle. We evaluate grasp on four relational benchmarks spanning molecular toxicity prediction (Tox21), mutagenesis, and citation matching (Cora), and show that it improves over purely symbolic, neural, and LLM-based baselines, while producing interpretable weighted rule ensembles. By replacing combinatorial search with gradient-guided LLM hypothesis generation, grasp retains boosting guarantees without sacrificing the transparency of symbolic outputs.
Figures & tables
Figure 1: grasp Framework. grasp frames relational learning as a neurosymbolic functional gradient boosting problem. At each iteration t , an LLM receives an example-centered relational context and proposes candidate symbolic rules. A logic solver evaluates each candidate against the relational data, assessing coverage and residual fit. The context updater computes pseudo-residuals ri=yi−σ(Ft−1(xi)) and selects high-residual examples to steer the LLM toward patterns not yet explained by the ensemble. The accepted rule φt is added with step size η , yielding Ft(x)=Ft−1(x)+ηφt(x) . After T iterations, grasp returns a weighted ensemble of interpretable probabilistic logic rules. This design replaces combinatorial symbolic search with gradient-guided LLM hypothesis generation, while retaining boosting convergence guarantees.
Dataset
BoostedRDN
NeuralLP
Grasp -Gem
Grasp -Cdx
One-shot
Citation and Mutagenesis domains
cora
95.02±1.96
91.86±0.00
92.99±0.51
97.67±0.26
74.78±21.57
mutagen
82.92±2.37
79.08±2.44
88.36±1.72
88.41±1.82
82.72±2.85
Tox21 domains
nr_ahr
81.98±0.47
68.80±0.07
76.11±1.45
81.41±0.65
63.81±7.73
nr_ar
75.84±1.36
78.82±0.19
74.81±1.71
78.11±3.04
61.73±10.80
Table 1: Predictive performance in terms of AUC-ROC across citation, mutagenesis, and Tox21 domains. Results are reported as percentages, with mean ± std over three trials. The best completed result for each dataset is shown in blue. Overall, GRASP variants are competitive with or outperform classical relational baselines across most domains, while the one-shot LLM baseline is less stable. “–” denotes no usable result.
Dataset
BoostedRDN
NeuralLP
Grasp -Gem
Grasp -Cdx
One-shot
Citation and Mutagenesis domains
cora
96.10±1.57
94.15±0.00
94.36±0.37
98.23±0.00
78.89±17.77
mutagen
86.53±1.47
89.53±1.48
94.36±0.65
93.65±0.77
90.23±2.51
Tox21 domains
nr_ahr
38.17±1.57
18.43±0.35
32.37±0.72
34.49±1.02
21.32±5.72
nr_ar
41.99±2.41
13.32±0.02
38.77±3.02
41.99±1.93
13.79±11.05
Table 2: Predictive performance in terms of AUC-PR across citation, mutagenesis, and Tox21 domains. Results are reported as percentages, with mean ± std over three trials. Since AUC-PR is more sensitive to class imbalance, these results highlight the ability of each method to recover positive examples in sparse relational prediction settings. The best result for each dataset is shown in blue bold. “–” denotes no usable result.
Table 3: Representative top-ranked rule per method on three targets (baselines first, Grasp last and highlighted with a plain-language reading). Grasp invents high-level predicates ( mol_num_rings , mol_count_fg ), uses numeric thresholds, and recovers chemically meaningful patterns compared to other methods. Top rule per method ( Grasp : ∣w1−w0∣ ; BoostedRDN : leaf value).
Figure 2: Evaluation of rule importance and ensemble efficiency via insertion (top) and deletion (bottom) curves, ordered by absolute discriminative weight. Gray dashed, and dotted lines denote random and reverse-ordering ablations for grasp (Codex). grasp models exhibit steep insertion asymptotes and rapid deletion degradation (outperforming baselines), confirming accurate weight calibration and demonstrating that a highly compact subset of rules captures the majority of predictive variance. Initial deletion robustness in grasp (Codex) reflects semantic redundancy among top-weighted rules.
Inductive Logic Programming (ILP) aims to learn interpretable first-order rules from data, but existing symbolic and neuro-symbolic approaches struggle to scale to noisy and probabilistic settings. Classical ILP relies on discrete combinatorial rule search and is brittle under uncertainty, while differentiable ILP methods typically depend on predefined rule templates or inaccurate fuzzy operators that suffer from vanishing gradients or poor approximation of logical structure when reasoning over probabilistic predicate valuations. This paper proposes an Attention-based Neuro-symbolic Differentiable Rule Extractor (ANDRE), a novel ILP framework that learns first-order logic programs by optimizing over a continuous rule space with attention-based logical operators. ANDRE replaces both rule templates and logical operators with fully differentiable, attention-driven conjunction and disjunction operators that approximate logical min-max semantics, enabling accurate, stable, and interpretable reasoning over probabilistic data. By softly selecting, negating, or excluding predicates within each rule, ANDRE supports flexible rule induction while preserving symbolic structure. Extensive experiments on classical ILP benchmarks, large-scale knowledge bases, and synthetic datasets with probabilistic predicates and noisy supervision demonstrate that ANDRE achieves competitive or superior predictive performance while reliably recovering correct symbolic rules under uncertainty. In particular, ANDRE remains robust to moderate label noise, substantially outperforming existing differentiable ILP methods in both rule extraction quality and stability.
Iman Sharifi, Peng Wei, Saber Fallah
Dept. of Mechanical and Aerospace Engineering, George Washington University, USA · Dept. of Mechanical Engineering Sciences, University of Surrey, UK
Large Language Models (LLMs) can propose rules in natural language, sidestepping the need for a predefined predicate space in traditional rule learning. Yet many LLM-based approaches ignore interactions among rules, and the opportunity to couple LLMs with probabilistic rule learning for robust inference remains underexplored. We present RLIE, a unified framework that integrates LLMs with probabilistic modeling to learn a set of weighted rules. RLIE has four stages: (1) Rule generation, where an LLM proposes and filters candidates; (2) Logistic regression, which learns probabilistic weights for global selection and calibration; (3) Iterative refinement, which updates the rule set using prediction errors; and (4) Evaluation, which compares the weighted rule set as a direct classifier with methods that inject rules into an LLM. We evaluate multiple inference strategies on real-world datasets. Applying rules directly with their learned weights yields superior performance, whereas prompting LLMs with the rules, weights, and logistic-model outputs surprisingly degrades accuracy. This supports the view that LLMs excel at semantic generation and interpretation but are less reliable for precise probabilistic integration. RLIE clarifies the potential and limitations of LLMs for inductive reasoning and couples them with classic probabilistic rule combination methods to enable more reliable neuro-symbolic reasoning.
Yang Yang, Hua XU, Zhangyi Hu +1
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511400, China · Institute of Deep Perception Technology, JITRI, Wuxi 214000, China
LLMs can solve program synthesis tasks but remain inefficient and unreliable on hard instances requiring large combinatorial search. Given a small set of reasoning traces, we use coding agents to compile them into reusable symbolic program synthesizers over constrained DSLs. The resulting solvers require no LLM calls at test time and are strong standalone systems: symbolic solver ensembles reach 91.3% accuracy on PBEBench-Lite and 84.7% on PBEBench-Hard, outperforming LLMs with test-time scaling for the latter by +16.3 percentage points at zero LLM inference cost. They also complement LLM search, improving PBEBench-Hard accuracy from 68.4% to 85.8% while reducing reported token usage by 78%, and raising SLR-Bench hard-tier accuracy from 34.4% to 58.0% in a neuro-symbolic hybrid setting. Compared to directly using coding agents as per-instance solvers, induced solvers are substantially more Pareto-efficient, amortizing a small one-time construction cost over many zero-token executions. Finally, most solvers transfer zero-shot to a real historical linguistics task - predicting sound changes in natural language data - reaching 80.1% accuracy under ensembling and recovering some plausible linguistic rules. Together, these results show that reasoning traces can be compiled into reusable symbolic solvers that solve many tasks directly, complement LLM inference on hard cases, and provide a scalable route to domain-general solver induction. We release code and data for reproducibility.
Atharva Naik, Yash Mathur, Prakam +2
Carnegie Mellon University · Independent Researcher