Probabilistic question-answering systems -- whether large language models (LLMs) themselves, retrieval-augmented generation (RAG), or trained multi-hop retrievers -- conflate "what is known" and "how to reason" into a single probabilistic computation: hallucination cannot be eradicated, evidence chains cannot be audited, and the system answers even when it does not know. We present LatWeave, which organizes knowledge into a multidimensional knowledge lattice and compiles multi-hop QA into three deterministic operators -- meet (constraint intersection), compare (lattice-order comparison), and abstain (structural abstention); LLMs appear only on the construction side (one-shot extraction) and the query-planning side, while the answer-generation path is zero-LLM, zero-task-training, and auditable end to end -- so that question answering over Web-published knowledge becomes reproducible item by item. Rather than claiming across-the-board SOTA, we characterize the operating envelope of this paradigm on six public benchmarks: when knowledge is complete (MetaQA, 39,093 questions) meet chains are near-lossless over three hops (any-hit 0.9975, on par with fully supervised KBQA); on templated multi-hop home ground (2WikiMultihopQA held-out n=1,258) EM 0.865, well above published structure-augmented RAG reproductions; on open-text deep composition (MuSiQue) and extraction-coverage gaps (HotpotQA) we report degradation honestly and attribute it to causes outside the lattice-algebra layer; and when information is incomplete (IIRC) we achieve structural abstention with abstain accuracy 0.971 and leak rate 0.029. Within the operating envelope, deterministic execution pays no performance penalty, and every step on the answer path can be recomputed -- precisely the source of end-to-end auditability.
Figures & tables
Figure 1 . LatWeave in two stages. Build time (offline, one-off): an LLM extracts typed predicates from question-blind documents (dedup, normalization and a dimension whitelist absorb extraction uncertainty before it can enter the lattice), yielding a product lattice whose instances each carry their source_text . Query time (online, per question): an LLM planner compiles the question into a typed plan, and a deterministic core executes the three operators — meet , compare , abstain . The LLM is confined to the two orange boxes (extraction and planning); the answer path carries zero LLM calls and zero task training, so every answer is an auditable meet chain.
Figure 2 . A product-lattice slice (species × habitat mini-example): (a) value hierarchies of the two dimensions; (b) an instance as a point in the product lattice; (c) one meet in three deterministic steps—a hit requires the descendant check on every dimension, and a conflict on any dimension entails abstain.
Table 4 . MetaQA test main results ( n=39,093 , single final run, frozen 9-shot configuration).
Method
1-hop
2-hop
3-hop
Paradigm (trained)
KV-Mem ( Miller et al., 2016 )
0.962
0.827
0.489
Memory networks (yes)
GraftNet ( Sun et al., 2018 )
0.970
0.948
0.777
Subgraph retrieval (yes)
SRN ( Qiu et al., 2020 )
0.970
0.951
0.752
Neural reasoning (yes)
PullNet ( Sun et al., 2019 )
0.970
0.999
0.914
Graph retrieval (yes)
EmbedKGQA ( Saxena et al., 2020 )
0.975
0.988
0.948
KGE (yes)
NSM ( He et al., 2021 )
0.971
0.999
0.989
State machine (yes)
Table 5 . MetaQA test vs. published KBQA methods (supervised rows: Hits@1 verified from source papers, converted to decimals). The Shrestha–Kim row reports Hit Rate—verbatim the any-hit measure, hence directly comparable to our any-hit row; their pooled micro-F1 is not cross-compared with our per-question set-F1 0.9508.
Figure 3 . Hop decay in two contrasts. MetaQA (complete KB, any-hit): lattice meet chains are nearly lossless over three hops (0.9990 / 0.9983 / 0.9955), while the same-corpus RAG baseline (hits@1) decays to 0.364 at 3-hop. MuSiQue (open-text extraction, EM): lattice retrieval collapses to zero on multi-hop text composition (0.048 / 0 / 0), while RAG stays flat at a low 0.12–0.14 (insensitive to hop count).
Level
Mechanism
Grounding object
Status
L0
Free generation
none
Rejected
L1
Global vocabulary
global relation names
Internalized (HotpotQA)
L2
Entity menu
seed entity’s lattice neighborhood
Internalized (FRAMES)
L3
Per-hop replanning
intermediate entities
Not adopted (argued below)
Table 6 . Four levels of query grounding.
Dataset
Knowledge-source form
Main bottleneck
Result
2WikiMultihopQA
Open text extraction, templated questions
—(home ground)
EM 0.865
MetaQA
KB triples mapped directly
Question parsing only
any-hit 0.9975
HotpotQA
Open text extraction, free paragraphs
Coverage gap (90.6%)
EM 0.1687
MuSiQue
Open text extraction, 4-hop chains
Chain deepening (48.3% hop-1 break)
EM 0.0248
FRAMES (dev n=742 )
Open text extraction, multi-hop + factuality
Coverage gap + chain-design error
EM 0.0256
IIRC (held-out n=130 )
Open text extraction, incomplete information
Answerability boundary
Prec. measure 0.9437
Table 7 . Operating envelope (all numbers are frozen final values).
Metric
Held-out 130
dev_tune 1,171
EM full (Phase-1-hit denominator)
0.0417 (1/24)
0.0264 (6/227)
Recall measure (answerable EM)
0.0105
0.0070
Resolved (Phase-1 hit rate)
25.3% (24/95)
26.4%
Leak rate (none answered)
0.029 (1/35)
0.016 (5/312)
Abstain accuracy
0.971 (34/35)
0.984
Precision measure (abstain × non-leak)
0.9437
0.9682
Table 8 . IIRC main results (held-out n=130 ; dev_tune as drift check).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Approach family
Zero-train
Determ.
Audit.
Abstain
Neural multi-hop retrievers
partial 1
×
×
×
RAG / GraphRAG
partial 2
×
eng.-level
×
KBQA / semantic parsing
× 3
✓
✓
×
Selective prediction
× 4
—
×
prob. thresh.
LatWeave (this work)
✓
✓
✓
✓
Appendix
Table 9 . Positioning: only LatWeave ticks all four columns. Zero-training is graded because several families mix trained and training-free members (see footnotes). Detection (auditability) is graded because chunk-ID provenance yields engineering-level auditability only: the step from retrieved passage to answer remains unverifiable.
Figure 4 . Query-time flow. The LLM planner compiles the question into a typed plan; deterministic execution on the lattice returns a meet hit, a compare verdict, or abstain ; the answer ships with its evidence text.
Figure 5 . Per-hop conditional meet accuracy P(correct∣chain reaches hop k) on the MuSiQue tune pool ( n=2,175 ), grouped by the question’s true depth.
Retrieval-Augmented Generation (RAG) has become a standard approach for knowledge-intensive question answering, but existing systems remain brittle on multi-hop questions, where solving the task requires chaining multiple retrieval and reasoning steps. Key challenges are that current methods represent reasoning through free-form natural language, where intermediate states are implicit, retrieval queries can drift from intended entities, and errors are detected by the same model that produces them making self-reflection an unreliable, ungrounded signal. We observe that multi-hop question answering is a typical form of step-by-step computation, and that this structured process aligns closely with how code-specialized language models are trained to operate. Motivated by this, we introduce \pyrag, a framework that reformulates multi-hop RAG as program synthesis and execution. Instead of free-form reasoning trajectories, \pyrag represents the reasoning process as an executable Python program over retrieval and QA tools, exposing intermediate states as variables, producing deterministic feedback through execution, and yielding an inspectable trace of the entire reasoning process. This formulation further enables compiler-grounded self-repair and execution-driven adaptive retrieval without any additional training. Experiments on five QA benchmarks (PopQA, HotpotQA, 2WikiMultihopQA, MuSiQue, and Bamboogle) show that \pyrag consistently outperforms strong baselines under both training-free and RL-trained settings, with especially large gains on compositional multi-hop datasets. Our code, data and models are publicly available at https://github.com/GasolSun36/PyRAG.
Jiashuo Sun, Jimeng Shi, Yixuan Xie +10
University of Illinois Urbana-Champaign · Hong Kong University of Science and Technology · Texas A&M University +1
Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA), which requires reasoning over evidence from multiple documents. Current multi-hop RAG methods generally focus on either query-side task decomposition or corpus-side knowledge graph construction. Despite their progress, these methods still struggle to achieve satisfactory performance on complex multi-hop QA tasks. To this end, we propose ConRAG, a consensus-driven multi-view RAG framework that effectively boosts LLMs on complex multi-hop QA. The core of ConRAG is to systematically optimize both the query and corpus sides and to leverage multi-view evidence (relation, entity, and text signals) for more accurate retrieval. Extensive experiments on three multi-hop QA benchmarks show that ConRAG consistently outperforms all baselines by a clear margin, e.g., up to +26.9% average performance gains over vanilla RAG, and enables Gemma-4-31B to achieve a new state-of-the-art record on the challenging MuSiQue benchmark.
Yikai Zhu, Kunfeng Chen, Qihuang Zhong +2
School of Computer Science, Wuhan University Wuhan, China
Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing. Despite these advances, LLMs and LLM-based systems remain prone to a variety of failure modes. Retrieval-augmented generation (RAG) systems have emerged as a common deployment scenario seeking to both avoid the well known risk of the LLM "hallucinating" information, and to enable reasoning and question answering over proprietary information that the LLM did not have access to during training without resorting to expensive model fine-tuning. In this work, we explore the idea of using a lightweight graph structure with a relatively simple graph schema, to support the RAG subsystem via a dedicated toolset. We design an agentic system with a variety of vector search and graph query tools operating over a structured dataset based on a curated subset of English Wikipedia articles, and evaluate its performance on questions from MoNaCo, a challenging Wikipedia QA benchmark of complex query answering tasks. Our results show that the introduction of graph-based tools can significantly increase the precision and recall of factual correctness, can halve the number of hallucinated answers, and achieves the highest fine-grained truthfulness score among the three evaluated scenarios. All this with a modest increase in token usage.
Christopher J. Wedge, Joshua Stutter, Danny Dixon +1
National Innovation Centre for Data, Newcastle University