The disambiguation of semantically similar statutory text across jurisdictions is a retrieval problem that existing methods do not solve. This inter-context conflict can steer generative models toward confidently produced answers grounded in topically relevant but jurisdictionally incorrect sources. Tobacco and nicotine regulations vary by US jurisdiction, often sharing similar language, thus, robust reasoning requires identifying which jurisdiction's law governs a given product, not merely retrieving relevant text. Emerging products (e.g., pouches) exploit ambiguous definitions to evade regulation. State-of-the-art (SOTA) document retrieval-augmented generation (RAG) methods struggle to address this inter-context conflict, and thus struggle to connect image attributes (e.g., rich attribute captions) to the set of similar legislation texts. We introduce NicoPRISM (Nicotine Product and Regulation Image-and-Text Surveillance Multimodal), comprising 161,563 images, attribute captions, a knowledge base of product, health, and legislative documents spanning 13 US jurisdictions, and 1,495 validated question-answer pairs across two tasks: policy compliance QA and product knowledge QA. We also propose PRISM-RAG, a multimodal hypergraph RAG framework built over images, captions, and entities without any LLM calls at index time, grounding every query in a product image and routes retrieval through a jurisdiction-aware context assembly mechanism guaranteeing that statutory text from the queried jurisdiction reaches the language model by construction. PRISM-RAG retrieves passages from the correct jurisdiction in 93.9% of policy compliance queries, a 48.6 percentage point advantage over standard RAG (p<0.001), using zero LLM calls at index time and one at query time, and is competitive with or outperforms SOTA RAG frameworks across keyword, semantic, jurisdiction-, and compliance-accuracy metrics.
Figures & tables
Figure 1: An image-caption sample from the NicoPRISM dataset of a nicotine pouch product with seven attribute captions. Best viewed in zoom and in color.
Figure 2: NicoPRISM Dataset statistics across the web set, TikTok set, and YouTube set, with image samples. Best viewed in zoom and in color.
Dataset
Images
Brands
Sources
Modality
Tasks
KB Types
Jurisdictions
QA Pairs
Expert Val.
Murthy et al. [ 1 ]
826
3
1
I
{ DET }
–
–
–
✗
Vassey et al. [ 2 ]
6,999
7
1
I
{ DET }
–
–
–
✗
PHAD [ 3 ]
171,900
8
2
I
{ DET }
–
–
–
✗
NicoPRISM (Ours)
161,563
277
3
I+T+D
{ RET , PKQ , PCQ }
Prod., Health, Legis.
13
1,495
✓
Table I: A comparison of NicoPRISM with existing tobacco and nicotine product datasets. Modality : I = image only; I+T+D = image, text captions, and documents. KB Types : document categories in the knowledge base (Prod. = product knowledge, Health = health context, Legis. = legislative texts). Tasks: DET = detection, RET = image retrieval, PKQ = product knowledge QA, PCQ = policy compliance QA. Expert Val. = expert-validated QA pairs.
Figure 3: Our proposed PRISM-RAG framework is organized into a large-scale hypergraph for multimodal RAG. (1) Concept hyperedges connect attributes and document entities. (2) Images to similar attributes form bridge hyperedges . (3) Entity-sentence hyperedges connect entities found within sentences within documents. (4) The assembled multimodal hypergraph combines nodes from steps (1–3) with no large model calls . Best viewed in color and with zoom.
Index Time
Query Time
Component
Bimodal Emb.
Concept HE
Bridge HE
Entity Extract.
Seed Retrieval
HE Scoring
Node Scoring s(q,v)
Context Assembly
Node sets
VI (image product)
✓
✓
✓
✗
✓
✓
✓
✗
VA (caption attribute)
✗
✓
✓
✗
✗
✓
✓
✗
VT (document entity)
✗
✓
✗
✓
✗
✓
✓
★
VJ (jurisdiction anchor)
✗
✗
✗
★
✗
✗
✗
★
Table II: Role of each node set and hyperedge type across all phases of PRISM-RAG. ✓ = participates directly; ★ = participates passively; ✗ = does not participate.
Method
KW-R
KW-F1
CC
JA
CA
Judge
RG
StandardRAG
0.2683
0.1821
0.3942
0.4495
0.3911
2.7442
0.7336
HippoRAG2
0.0507
0.0345
0.5778
0.1408
0.0844
1.4106
0.4725
HyperGraphRAG †
0.2220
0.1475
0.2816
0.8773
0.0133
2.7283
0.8850
PRISM-RAG (ours)
0.3139
0.1997
0.5495
0.9387
0.4178
2.8551
0.7287
Table III: Policy compliance QA results on 1,325 test instances spanning 13 jurisdictions. KW-R = keyword recall; KW-F1 = keyword F1; CC = context coverage; JA = retrieval-level jurisdiction accuracy; CA = compliance accuracy; Judge = LLM-as-judge score (1–5); RG = response groundedness. All PRISM-RAG results use m=10 , c=50 . † HyperGraphRAG makes two LLM calls and three API calls total per query. Bootstrap 95% CIs and BH-corrected p -values for all pairwise comparisons are reported in the Appendix.
Comparison
JA [95% CI]
CA [95% CI]
KW-F1 [95% CI]
Judge [95% CI]
vs. StandardRAG
+0.486[0.462,0.509]∗
+0.030\ [{-0.034},\ 0.094]\
+0.019[0.014,0.023]∗
+0.114[0.043,0.180]∗
vs. HippoRAG2
+0.796[0.781,0.811]∗
+0.340[0.281,0.404]∗
+0.166[0.159,0.173]∗
+1.455[1.381,1.528]∗
vs. HyperGraphRAG
+0.061[0.041,0.081]∗
+0.409[0.345,0.472]∗
+0.053[0.049,0.058]∗
+0.136[0.049,0.219]∗
Table IV: Pairwise mean gaps (PRISM-RAG minus baseline) with bootstrap 95% confidence intervals and BH-corrected p -values on policy compliance QA. The star ∗ means a significant result at adjusted q<0.05 (BH correction over all 24 simultaneously tested hypotheses across metrics and baseline pairs; 10,000 bootstrap resamples and 10,000 permutations). CA uses N=235 matched instances with an extractable compliance label; all other metrics use N=1,387 matched pairs ( N=1,231 for PRISM-RAG vs. HyperGraphRAG on JA).
Method
Index LLM
Query LLM
WT (s/q)
Avg. Tok
StandardRAG
0
1
1.99
3970
HippoRAG2
∣D∣
1
24.15
366310
HyperGraphRAG
∣D∣
2
10.88
1438
PRISM-RAG (ours)
0
1
3.37
4331
Table V: Index-time and query-time cost across RAG methods. Index LLM = LLM calls at index time; Query LLM = LLM calls per query; WT = mean wall time per query (s/q); Avg. Tok = mean input tokens per LLM call with respect to gpt-5.4-nano . ∣D∣=2,001 document chunks from NicoPRISM’s document corpus.
Method
KW-P
KW-R
KW-F1
CC
JA
CA
Judge
Index-time: bridge method and entity extractor
PRISM-RAG (ours)
0.1626
0.3139
0.1997
0.5495
0.9387
0.4178
2.8551
PRISM-RAG bridge
0.1642
0.3140
0.2008
0.5565
0.9396
0.4000
2.8891
PRISM-RAG LLM
0.1611
0.3147
0.1985
0.5629
0.8946
0.4489
2.8415
PRISM-RAG ontology
0.1639
0.3125
0.2000
0.5548
0.9094
0.3778
2.8838
PRISM-RAG agglo
0.1618
0.3061
0.1970
0.5342
0.8682
0.4356
2.8581
Table VI: Policy compliance QA ablation results. CC = context coverage; JA = retrieval-level jurisdiction accuracy; CA = compliance accuracy; Judge = LLM-as-judge score (1–5). JA is the primary ablation metric; CA is reported alongside it but exhibits a JA-CA dissociation in several configurations (see text). Unless otherwise noted, all ablations use c=50 , m=10 .
Method
KW-P
KW-R
KW-F1
CC
Main comparison (from Table III )
StandardRAG
0.1518
0.1935
0.1360
0.1520
HippoRAG2
0.0326
0.0137
0.0173
0.3395
HyperGraphRAG
0.1189
0.1577
0.1069
0.0473
PRISM-RAG (ours)
0.1322
0.2027
0.1342
0.1552
Index-time: bridge method and entity extractor
Table VII: Product knowledge QA ablation results. CC = context coverage. Performance is broadly uniform across ablations, consistent with this task being bounded by knowledge-base coverage rather than retrieval architecture. Unless otherwise noted, all ablations use c=50 , m=10 .
Figure 4: Qualitative comparison of PRISM-RAG and StandardRAG on a Washington, D.C. policy compliance question. Green highlights mark ground-truth keywords that appear in the response (contributing to KW-R); orange highlights mark DC statutory citations (§ 7-1721.01(1), § 7-1721.08(b)) that are correct but excluded from KW-R computation due to the ‘§’ in the keyword string, understating PRISM-RAG’s factual accuracy on this instance. PRISM-RAG retrieves DC-specific statutory text via its jurisdiction-aware direct injection path (JA = 1.0 for this instance), enabling it to cite the applicable DC flavored tobacco product framework verbatim. StandardRAG retrieves semantically similar federal FDA regulatory content, which is not relevant for the question scope.
Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority structures. Unlike traditional multi-hop or legal QA, this task requires structured procedural lookups and evidence-set closure rather than entity resolution or case-law reasoning. Existing RAG systems struggle here due to flattened citation edges, fragmented retrieval expansions, and fragile post-hoc attribution. We formalize Regulatory Compliance QA with RegOps-Bench, a novel benchmark featuring an Operational Knowledge Graph derived from complex national R&D regulations. To address these bottlenecks, we propose RefWalk, a unified framework driven by a shared topic anchor. RefWalk traverses cross-document citations, fuses multi-view candidates via max-based aggregation, and enforces per-rule attribution to explicitly map claims to sources. We establish a strong baseline with substantial improvements in retrieval recall and citation accuracy. Finally, a contrastive evaluation on a U.S. health compliance dataset (HIPAA) reveals that existing systems exhibit saturation on flat-structure rules, underscoring the need for RegOps-Bench. Our code is available at https://github.com/yeongjoonJu/RefWalk.
Yeong-Joon Ju, Seong-Whan Lee
Department of Artificial Intelligence, Korea University
Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. However, its reasoning process remains largely opaque: intermediate reasoning steps are difficult to verify and cannot be reliably attributed to specific evidence. Moreover, missing user-specific context is rarely detected systematically, often leading to incomplete or incorrect output. We propose NeSy-RAG, a modular neuro-symbolic RAG framework that synthesizes attributable Prolog modules from retrieved text chunks. For each chunk, the system generates semantically meaningful predicates that encode Boolean claims, which may depend on user facts. Using joint natural language-code embeddings, predicates are retrieved and composed into Prolog queries. To address incomplete user context, we introduce a symbolic knowledge-gap detection mechanism that identifies missing user facts whose truth values affect the query outcome and automatically triggers follow-up interactions. Executing the resulting Prolog queries yields deterministic answers together with transparent execution traces that link each reasoning step to its originating source. On the ShARC benchmark, without domain-specific training, NeSy-RAG achieves 61.1% accuracy, outperforming a same-model RAG baseline that achieves 42.8% accuracy.
Jonas Gann, Michael Gertz
Data Science Group, Heidelberg University, Germany
Retrieval-Augmented Generation (RAG) has become a standard approach for knowledge-intensive question answering, but existing systems remain brittle on multi-hop questions, where solving the task requires chaining multiple retrieval and reasoning steps. Key challenges are that current methods represent reasoning through free-form natural language, where intermediate states are implicit, retrieval queries can drift from intended entities, and errors are detected by the same model that produces them making self-reflection an unreliable, ungrounded signal. We observe that multi-hop question answering is a typical form of step-by-step computation, and that this structured process aligns closely with how code-specialized language models are trained to operate. Motivated by this, we introduce \pyrag, a framework that reformulates multi-hop RAG as program synthesis and execution. Instead of free-form reasoning trajectories, \pyrag represents the reasoning process as an executable Python program over retrieval and QA tools, exposing intermediate states as variables, producing deterministic feedback through execution, and yielding an inspectable trace of the entire reasoning process. This formulation further enables compiler-grounded self-repair and execution-driven adaptive retrieval without any additional training. Experiments on five QA benchmarks (PopQA, HotpotQA, 2WikiMultihopQA, MuSiQue, and Bamboogle) show that \pyrag consistently outperforms strong baselines under both training-free and RL-trained settings, with especially large gains on compositional multi-hop datasets. Our code, data and models are publicly available at https://github.com/GasolSun36/PyRAG.
Jiashuo Sun, Jimeng Shi, Yixuan Xie +10
University of Illinois Urbana-Champaign · Hong Kong University of Science and Technology · Texas A&M University +1