Meet, Compare, or Abstain: LatWeave for Deterministic Multi-Hop Question Answering on Knowledge Lattices
Organizations: ZenSmart Technology (Beijing) Co., Ltd., China
Abstract
Probabilistic question-answering systems -- whether large language models (LLMs) themselves, retrieval-augmented generation (RAG), or trained multi-hop retrievers -- conflate "what is known" and "how to reason" into a single probabilistic computation: hallucination cannot be eradicated, evidence chains cannot be audited, and the system answers even when it does not know. We present LatWeave, which organizes knowledge into a multidimensional knowledge lattice and compiles multi-hop QA into three deterministic operators -- meet (constraint intersection), compare (lattice-order comparison), and abstain (structural abstention); LLMs appear only on the construction side (one-shot extraction) and the query-planning side, while the answer-generation path is zero-LLM, zero-task-training, and auditable end to end -- so that question answering over Web-published knowledge becomes reproducible item by item. Rather than claiming across-the-board SOTA, we characterize the operating envelope of this paradigm on six public benchmarks: when knowledge is complete (MetaQA, 39,093 questions) meet chains are near-lossless over three hops (any-hit 0.9975, on par with fully supervised KBQA); on templated multi-hop home ground (2WikiMultihopQA held-out n=1,258) EM 0.865, well above published structure-augmented RAG reproductions; on open-text deep composition (MuSiQue) and extraction-coverage gaps (HotpotQA) we report degradation honestly and attribute it to causes outside the lattice-algebra layer; and when information is incomplete (IIRC) we achieve structural abstention with abstain accuracy 0.971 and leak rate 0.029. Within the operating envelope, deterministic execution pays no performance penalty, and every step on the answer path can be recomputed -- precisely the source of end-to-end auditability.
Figures & tables
| Dataset | Size used | Structure | Role |
|---|---|---|---|
| MetaQA ( Zhang et al., 2018 ) | test | 1/2/3-hop: 9,947 / 14,872 / 14,274 | Control: complete knowledge lattice algebra lossless |
| 2WikiMultihopQA ( Ho et al., 2020 ) | held-out | 4 question types | Main experiment: templated multi-hop, end-to-end |
| HotpotQA ( Yang et al., 2018 ) | held-out | bridge 592 / comparison 149 | Boundary 1: fact-coverage gap |
| MuSiQue ( Trivedi et al., 2022 ) | held-out | 2/3/4-hop: 125/76/41 | Boundary 2: longer-chain degradation |
| IIRC ( Ferguson et al., 2020 ) | held-out | span 59/ value 23/ binary 13/ none 35 | Refusal experiment: abstain under incompleteness |
| FRAMES ( Krishna et al., 2025 ) | dev 742 + held-out 82 | factuality multi-hop | Compounded stress test: coverage grounding |
| Method | EM | F1* | Resolved |
| Keyword baseline | 0.083 | 0.111 | 73.2% |
| Naive RAG (TF-IDF + LLM) | 0.198 | 0.205 | 100% |
| GraphRAG ( Edge et al., 2024 ) | 0.514 ‡ | 0.586 ‡ | — |
| HippoRAG 2 ( Gutiérrez et al., 2025 ) | 0.650 ‡ | 0.710 ‡ | — |
| IRCoT ( Trivedi et al., 2023 ) | 0.53 | 0.65 | — |
| HopRAG ( Liu et al., 2025 ) | 0.62 | 0.69 | — |
| Question type | Share | EM | vs. overall |
|---|---|---|---|
| bridge_comparison | 21.9% | 0.975 | +11.0pp |
| comparison | 24.2% | 0.888 | +2.3pp |
| compositional | 41.7% | 0.815 | 5.0pp |
| inference | 12.3% | 0.794 | 7.1pp |
| Overall | 100% | 0.865 | — |
| Hop | any-hit | recall (complete) | EM (strict) | set-F1 | |
|---|---|---|---|---|---|
| 1-hop | 9,947 | 0.9990 | 0.9987 | 0.9447 | 0.9798 |
| 2-hop | 14,872 | 0.9983 | 0.9983 | 0.6239 | 0.9448 |
| 3-hop | 14,274 | 0.9955 | 0.9954 | 0.5304 | 0.9369 |
| All | 39,093 | 0.9975 | 0.9973 | 0.6714 | 0.9508 |
| Method | 1-hop | 2-hop | 3-hop | Paradigm (trained) |
| KV-Mem ( Miller et al., 2016 ) | 0.962 | 0.827 | 0.489 | Memory networks (yes) |
| GraftNet ( Sun et al., 2018 ) | 0.970 | 0.948 | 0.777 | Subgraph retrieval (yes) |
| SRN ( Qiu et al., 2020 ) | 0.970 | 0.951 | 0.752 | Neural reasoning (yes) |
| PullNet ( Sun et al., 2019 ) | 0.970 | 0.999 | 0.914 | Graph retrieval (yes) |
| EmbedKGQA ( Saxena et al., 2020 ) | 0.975 | 0.988 | 0.948 | KGE (yes) |
| NSM ( He et al., 2021 ) | 0.971 | 0.999 | 0.989 | State machine (yes) |
| Level | Mechanism | Grounding object | Status |
|---|---|---|---|
| L0 | Free generation | none | Rejected |
| L1 | Global vocabulary | global relation names | Internalized (HotpotQA) |
| L2 | Entity menu | seed entity’s lattice neighborhood | Internalized (FRAMES) |
| L3 | Per-hop replanning | intermediate entities | Not adopted (argued below) |
| Dataset | Knowledge-source form | Main bottleneck | Result |
|---|---|---|---|
| 2WikiMultihopQA | Open text extraction, templated questions | —(home ground) | EM 0.865 |
| MetaQA | KB triples mapped directly | Question parsing only | any-hit 0.9975 |
| HotpotQA | Open text extraction, free paragraphs | Coverage gap (90.6%) | EM 0.1687 |
| MuSiQue | Open text extraction, 4-hop chains | Chain deepening (48.3% hop-1 break) | EM 0.0248 |
| FRAMES (dev ) | Open text extraction, multi-hop factuality | Coverage gap chain-design error | EM 0.0256 |
| IIRC (held-out ) | Open text extraction, incomplete information | Answerability boundary | Prec. measure 0.9437 |
| Metric | Held-out 130 | dev_tune 1,171 |
|---|---|---|
| EM full (Phase-1-hit denominator) | 0.0417 (1/24) | 0.0264 (6/227) |
| Recall measure (answerable EM) | 0.0105 | 0.0070 |
| Resolved (Phase-1 hit rate) | 25.3% (24/95) | 26.4% |
| Leak rate (none answered) | 0.029 (1/35) | 0.016 (5/312) |
| Abstain accuracy | 0.971 (34/35) | 0.984 |
| Precision measure (abstain non-leak) | 0.9437 | 0.9682 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Approach family | Zero-train | Determ. | Audit. | Abstain |
|---|---|---|---|---|
| Neural multi-hop retrievers | partial 1 | |||
| RAG / GraphRAG | partial 2 | eng.-level | ||
| KBQA / semantic parsing | 3 | ✓ | ✓ | |
| Selective prediction | 4 | — | prob. thresh. | |
| LatWeave (this work) | ✓ | ✓ | ✓ | ✓ |
| Configuration | EM | Resolved |
|---|---|---|
| Full system | 0.873 | 92.0% |
| train gold transfer (dev-only) | 0.731 | 85.3% |
| alias normalization | 0.853 | 90.8% |
| rule-based re-extraction (E15) | 0.843 | 90.4% |
| any-hit | recall (complete) | EM (strict) | ||
|---|---|---|---|---|
| 1 | 14,834 | 0.9974 | 0.9974 | 0.8332 |
| 2 | 5,729 | 0.9977 | 0.9977 | 0.7413 |
| 3–5 | 7,277 | 0.9974 | 0.9968 | 0.6442 |
| 6–10 | 4,041 | 0.9968 | 0.9968 | 0.5494 |
| 11 | 7,212 | 0.9978 | 0.9976 | 0.3788 |
| All | 39,093 | 0.9975 | 0.9973 | 0.6714 |
| Dataset | Attribution | Count (share) |
| HotpotQA | Extraction coverage gap | 667/736 (90.6%) |
| (residual broken chains) | H1 missing / H2 missing / seed no-out-edge | 527 / 81 / 59 |
| Query-layer planner | 69/736 (9.4%) | |
| MuSiQue | Break at hop 1 | 117/242 (48.3%) |
| (chain-break locus) | Break at hop 2 | 81/242 (33.5%) |
| Break at hop 3 / hop 4 | 16 / 4 (6.6% / 1.7%) |