Graph agents extend large language models (LLMs) with the ability to actively explore and reason over knowledge graphs through multi-step interactions with graph tools. However, training capable graph agents typically requires large collections of question-answer pairs and reasoning trajectories, whose manual construction is costly and difficult to scale. Moreover, employing proprietary LLMs to generate such supervision further risks exposing sensitive graph data to external services. Therefore, we propose GraphCert to bootstrap agentic graph reasoning with certified evidence rubrics during post-training. Specifically, the Bootstrapped Graph Quizzer guided by generation controls produces graph-grounded QA pairs and marks supporting evidence, which undergo execution certification and semantic curation. The accepted evidence is then canonicalized into certified evidence rubrics that later reward Graph Solver evidence alignment alongside answer correctness during GRPO training. Experiments on five graph reasoning domains in GRBENCH demonstrate that GraphCert consistently outperforms substantially larger LLM agents and post-training method. Furthermore, our analysis demonstrates that the learned policy transfers robustly across heterogeneous graph domains, suggesting that GraphCert acquires reusable graph-reasoning capabilities rather than domain-specific patterns. These results establish executable self-certification as an effective approach to self-training compact graph reasoning agents. Our code will be made publicly available.
Figures & tables
Figure 1: Overview of GraphCert. A local Quizzer generates graph-grounded QA candidates and supporting evidence, which are certified and converted into evidence rubrics. These rubrics provide hidden structural supervision for evidence-aware Solver post-training.
Healthcare
Literature
Academic
E-Commerce
Legal
Method
LLM
QwenScore
F1
QwenScore
F1
QwenScore
F1
QwenScore
F1
QwenScore
F1
BaseLLM
GPT-4o
0.137
0.048
0.221
0.064
0.097
0.080
0.110
0.080
0.244
0.110
DeepSeek-V3.2
0.104
0.075
0.192
0.111
0.134
0.140
0.095
0.115
0.222
0.265
TextRAG
GPT-4o
0.074
0.059
0.179
0.116
0.098
0.090
0.200
0.181
0.256
0.232
DeepSeek-V3.2
0.085
0.060
0.196
0.137
0.146
0.158
0.115
0.138
0.333
0.365
GraphRAG
GPT-4o
0.156
0.129
0.217
0.136
0.105
0.092
0.315
0.308
0.239
0.233
Table 1: Performance Comparison on GRBENCH Across Baseline Methods and GraphCert. Best results are in bold, and second-best results are underlined.
Statistic
Healthcare
Literature
Candidates
3,126
3,096
Accepted strict (%)
34.87
29.40
Accepted partial (%)
29.10
35.20
Rejected (%)
36.03
35.40
Certified examples
2,000
2,000
Table 2: Pipeline status before training. Strict and partial examples enter Dcert only after semantic curation. Rejected examples combine program and semantic failures.
N
Answer Supported
Evidence Sufficient
Query Consistent
Strict Grounding
Plain
300
38.7
38.7
45.3
36.7
GraphCert
300
82.7
72.7
77.3
71.0
Table 3: Independent assessment of generated QA quality (%). The blinded LLM judge is used for relative comparison rather than as a calibrated ground-truth oracle. Plain denotes the quizzer without certification or semantic curation.
Variant
Healthcare
Literature
GraphCert
0.722
0.681
w/o Certification
0.612
0.624
w/o Sevidence
0.566
0.654
Table 4: F1 results of controlled ablation experiments on the Healthcare and Literature domains. All variants use the same training data size and optimization budget.
Method
Literature
E-Commerce
GraphCert(Untrained)
0.301
0.212
GraphCert
0.681
0.583
GraphScout(DS)
0.646
0.562
GraphScout(4B)
0.618
0.481
Table 5: Quizzer-scale comparison reported in terms of F1 metric. GraphScout(DS) denotes the original GraphScout which uses DeepSeek-V3.2 as the Quizzer
Test Domain
Training Domain
Health.
Lit.
Acad.
E-Com.
Legal
Healthcare
0.722
0.544
0.500
0.496
0.637
Literature
0.672
0.681
0.537
0.602
0.644
Academic
0.561
0.587
0.674
0.571
0.655
E-Commerce
0.614
0.656
0.488
0.583
0.624
Legal
0.643
0.691
0.648
0.571
0.648
Table 6: Cross-domain transfer in terms of F1 metric. Each row specifies the training domain and each column specifies the test domain.
Backbone
Healthcare
Literature
Qwen3-4B-Instruct-2507
0.744/0.722
0.656/0.681
Qwen3-8B
0.609/0.601
0.663/0.687
Table 7: Scaling results reported as QwenScore/F1. Both rows use the full certified self-training method.
Table 8: Quizzer generation controls and their compatibility constraints. The five controls specify answer form, abstract question structure, graph-access structure, return semantics, and intended complexity; all are soft generation targets rather than verified labels.
Hyperparameter
Value
RL algorithm
GRPO
Training framework
verl
Training samples per domain
2,000
Optimization steps
400
Responses per prompt
8
Rollout temperature
1.0
Appendix
Table 9: GRPO hyperparameters used for GraphCert.
Configuration
4B experiments
8B experiments
CPU
Intel Xeon Gold 6342, 2.80 GHz
Intel Xeon Platinum 8358P, 2.60 GHz
GPU
4× NVIDIA A40
4× NVIDIA A800-SXM4-80GB
System memory
1 TB
1 TB
CUDA
12.9
12.3
Operating system
Ubuntu 22.04.5 LTS
Ubuntu 22.04.3 LTS
Appendix
Table 10: Hardware and software configurations used for the 4B and 8B experiments.
Figure 2: Average token consumption on Healthcare and Literature; lower is better. The horizontal axis uses a logarithmic scale.
Figure 3: QwenScore by question difficulty on the informative Easy and Medium splits of Healthcare and Literature.
Figure 4: Match-family distributions of semantically accepted questions in the biomedical run, comparing strict certificates with the combined strict-and-partial pool.
Figure 5: Question-pattern distributions of semantically accepted questions in the biomedical run, comparing strict certificates with the combined strict-and-partial pool.
LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit. In particular, deterministic computations over raw data and semantic deductions over natural-language claims are often entangled in an unstructured stream, leaving numerical conclusions hard to reproduce and qualitative judgments hard to inspect. To address this, we propose VeriGraph, a traceable neuro-symbolic reasoning framework that enables agents to construct an explicit heterogeneous evidence directed acyclic graph (DAG) during execution. VeriGraph introduces three evidence-expansion primitives, namely computational, grounding, and derivational expansion, to connect raw data, interpreter variables, computed results, and natural-language claims in a unified graph. Under this formulation, structural traceability is reduced to graph reachability from raw data sources to terminal claims, while semantic support is measured by claim-level evidence evaluation. To improve graph construction, we further design a graph-based policy optimization strategy with a composite reward that jointly supervises answer correctness, computational integrity, and derivational coherence. Experiments on four benchmarks show that VeriGraph-8B achieves the highest overall score among all baselines. More importantly, VeriGraph produces auditable evidence graphs with substantially stronger claim grounding, achieving a 87.61% Grounding Rate under our claim-level evidence support evaluation. These results suggest that explicit evidence-graph construction is a promising path toward verifiable data-analytic agents. Our code is available at https://github.com/ignorejjj/VeriGraph.
Jiajie Jin, Zhao Yang, Wenle Liao +5
Gaoling School of Artificial Intelligence, Renmin University of China
Graphs model relational data throughout science and industry, from citation networks to product co-purchase graphs. Because the nodes of many such graphs carry rich text, a growing line of work applies large language models (LLMs) to graph analysis. The most graph-native of these methods use graph tokens: a graph encoder compresses a graph view, such as a node, its k-hop neighbourhood, or a cluster, into a short block of continuous tokens that jointly encodes node attributes and topology and is read directly by the model. Existing methods, however, use graph tokens in a static single-shot manner: they encode one predefined graph view before the model has even seen the target and never revise it, leaving the model's step-by-step reasoning ability unused. We introduce agentic graph token reasoning, which recasts graph tokenization as part of the reasoning process itself. At each step, the model chooses which graph view to encode and at what granularity; a graph encoder is invoked on demand to materialise the corresponding graph tokens; and the resulting block is spliced into the running context. The model thus reasons step by step in the graph token space, and the tokens it reads are trajectory-dependent. We realise this with a three-stage training pipeline: (i) self-supervised tasks that teach the model to read heterogeneous graph tokens, (ii) a token-robust trajectory stage with a graph-token consistency regulariser, and (iii) preference optimisation that rewards trajectories in which the graph-token evidence and the node-text evidence agree. Across evaluations spanning seven graph domains, our models outperform a broad set of baselines by a large margin and transfer zero-shot to unseen domains without any per-target fine-tuning. More broadly, this work pushes LLM-based graph analysis from static graph-token encoders towards a graph-native agent paradigm.
Zhuoyi Peng, Yi Yang
The Hong Kong University of Science and Technology
Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems in 44 task-structure settings, with over 100,000 examples across four representations: natural language, structured language, adjacency list, and adjacency matrix. Evaluating eight LLMs on GT Bench shows that accuracy is strongly tied to the input representation, that the best representation shifts with graph density, size, and topology as well as with the model, and that this sensitivity persists, attenuated, in the strongest reasoning models. Building on these observations, we propose the Graph Theory Agent (GTA), which pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM. GTA lifts Phi-4 from 53.5% to 69.1% on the benchmark's easy split and from 33.0% to 41.5% on its hard split, outperforming eight prompting and agent baselines, and transfers without retraining to GraCoRe and NLGraph. Code for benchmark generation and evaluation: https://github.com/xzx34/GTA. The project homepage is available at https://xzx34.github.io/gta/.
Zixiang Xu, Yanbo Wang, Chenxi Wang +6
University of Southern California · University of California, Los Angeles · Mohamed bin Zayed University of Artificial Intelligence +2