Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis
Organizations: Computational Social Science, ETH Zurich, Zurich, Switzerland · Complexity Science Hub, Vienna, Austria
Abstract
Public commercial registries are formally open, yet their practical analysis remains difficult because relevant facts are scattered across millions of records that combine structured metadata, multilingual legal notices, temporal events, and entity aliases. This paper presents a controlled, tool-mediated agentic GraphRAG architecture for auditable natural-language analysis of such registries. The proposed pipeline transforms publications from the Swiss Official Gazette of Commerce into a Neo4j knowledge graph comprising over five million nodes and 4.7 million relationships. It combines deterministic ingestion of structured registry fields, LLM-assisted extraction of latent actors from unstructured notices, and a deterministic identity-resolution layer. An analytical agent operates on this graph through intent routing, restricted graph tools, bounded reflection, and state-machine-guided response synthesis. We evaluate the system using a multi-tier protocol covering answer quality, retrieval behavior, entity resolution, and multi-turn conversational performance. The complete architecture is compared with dense, lexical, and hybrid flat-retrieval baselines and with controlled architectural ablations. On a manually curated benchmark, graph-mediated retrieval increases factual correctness from 0.26 for the strongest flat-retrieval baseline to 0.83 for the complete system, with comparable improvements in relevance and completeness. Ablation results show that bounded reflection improves answer quality while intent routing and LLM-based graph enrichment improve reliability in difficult entity resolution tasks. An exploratory dashboard displays the graph evidence and execution traces underlying each response, allowing users to inspect how answers were produced.
Figures & tables
| Architecture Phase | Calls | Purpose | Configuration |
|---|---|---|---|
| 1. Intent Router | 1 | Classifies user intent to dynamically restrict the available sandbox tools. | GPT-4o mini (2024-07-18) |
| 2. Reflection Loop | 0–4 | Autonomously structures JSON tool requests and interprets raw graph responses. | GPT-4o mini (2024-07-18) , tool calling |
| 3. Response Synthesis | 1 | Formats the accumulated evidence according to UI constraints (e.g., Markdown tables or dossiers). | GPT-4o mini (2024-07-18) , tool calling |
| Note: Total calls per user query range from 2 to 6, bounded by the reflection-loop limit of four. | |||
| Dataset Name | Size | Generation Method | Evaluation Focus |
|---|---|---|---|
| Graph-Seeded Automated Benchmark | 300 questions | Graph-seeded generation | Agent trajectory (Tier 2), semantic quality (Tier 3) |
| Manually Curated Benchmark | 60 questions | Author-curated | Mean Correctness score (Tier 3) |
| Conversational Benchmark | 10 evaluated conversations / 36 turns | Author-curated | Context carryover and multi-turn coherence (Tier 4) |
| Metric | Result |
|---|---|
| NameHub Nodes Sampled | 1,000 |
| Grouped-Name Comparisons Evaluated | 1,191 |
| Orthographic Consistency Rate | 96.22% |
| Extracted item | Precision | Recall | F1 | Matching criterion |
|---|---|---|---|---|
| Persons | 0.989 | 0.956 | 0.972 | Fuzzy name match |
| Companies | 0.933 | 0.651 | 0.767 | Fuzzy name match, excluding structured subject entities |
| Roles | 0.723 | 0.611 | 0.662 | Coarse-compatible role match |
| Entity-role relations | 0.701 | 0.482 | 0.571 | Coarse-compatible entity-role match |
| Metric | Value |
|---|---|
| Questions Evaluated | 300 |
| Search-First Accuracy | 55.7% |
| Fallback Activation Rate | 2.0% |
| Average Reasoning Steps | 1.7 |
| Query Success Rate | 100% |
| Average Latency | 12.53 s |
| Architecture variant | Faith. | Ans. rel. | Info. recall | Mean s | Unsucc. |
|---|---|---|---|---|---|
| Dense Vector-RAG | 0.905 | 0.081 | 0.066 | 32.45 | 0.000 |
| Lexical Full-Text RAG | 0.975 | 0.047 | 0.030 | 0.88 | 0.940 |
| Hybrid Dense+Lexical RAG | 0.879 | 0.127 | 0.091 | 1.35 | 0.813 |
| GraphRAG w/o reflection | 0.894 | 0.425 | 0.388 | 5.54 | 0.053 |
| GraphRAG w/o router | 0.887 | 0.742 | 0.581 | 9.25 | 0.280 |
| Structured-only GraphRAG | 0.865 | 0.651 | 0.534 | 9.85 | 0.330 |
| Architecture variant | Correctness | Ans. rel. | Info. recall | Mean s |
|---|---|---|---|---|
| Dense Vector-RAG | 0.143 | 0.247 | 0.118 | 15.62 |
| Lexical Full-Text RAG | 0.237 | 0.322 | 0.261 | 14.29 |
| Hybrid Dense+Lexical RAG | 0.262 | 0.470 | 0.237 | 17.50 |
| GraphRAG w/o reflection | 0.552 | 0.563 | 0.560 | 9.57 |
| GraphRAG w/o router | 0.708 | 0.777 | 0.703 | 10.72 |
| Structured-only GraphRAG | 0.545 | 0.733 | 0.578 | 16.63 |
| Architecture variant | Evidence-bearing | Supported outcome | Empty result | Tool/error | Retrieved documents / tool calls |
|---|---|---|---|---|---|
| Dense Vector-RAG | 0.283 | 0.067 | 0.000 | 0.000 | 16.32 |
| Lexical Full-Text RAG | 0.300 | 0.133 | 0.667 | 0.000 | 4.62 |
| Hybrid Dense+Lexical RAG | 0.533 | 0.167 | 0.000 | 0.000 | 20.00 |
| GraphRAG w/o reflection | 0.583 | 0.550 | 0.133 | 0.150 | 1.00 |
| GraphRAG w/o router | 0.850 | 0.683 | 0.000 | 0.117 | 1.57 |
| Structured-only GraphRAG | 0.783 | 0.483 | 0.000 | 0.450 | 2.07 |
| Architecture variant | Corr. | Ans. rel. | Info. recall | Turn succ. | Carryover | Tool trans. |
|---|---|---|---|---|---|---|
| Dense Vector-RAG | 0.344 | 0.644 | 0.317 | 0.367 | 0.433 | N/A |
| Lexical Full-Text RAG | 0.342 | 0.493 | 0.345 | 0.308 | 0.400 | N/A |
| Hybrid Dense+Lexical RAG | 0.310 | 0.677 | 0.283 | 0.308 | 0.350 | N/A |
| GraphRAG w/o reflection | 0.458 | 0.598 | 0.450 | 0.450 | 0.517 | 0.800 |
| GraphRAG w/o router | 0.554 | 0.714 | 0.525 | 0.542 | 0.517 | 0.900 |
| Structured-only GraphRAG | 0.604 | 0.827 | 0.580 | 0.608 | 0.650 | 0.900 |
| Extracted Role | Graph Edge Type |
|---|---|
| SUBJECT | HAS_EVENT |
| PARENT | HEAD_OFFICE_OF |
| SELLER | TRANSFERRED_TO |
| BUYER | ACQUIRED_FROM |
| DEBTOR_COMPANY | HAS_EVENT |
| DEBTOR_PERSON | INVOLVED_IN |
| Node Type | Strong | Weak | Total |
|---|---|---|---|
| Company | 935,014 | 112,344 | 1,047,358 |
| Person | 57,508 | 606,392 | 663,900 |
| Event | N/A | N/A | 2,180,645 |
| NameHub | N/A | N/A | 1,354,406 |
| Total Nodes | 5,246,309 | ||
| Total Edges | 4,729,070 |
| Semantic Intent (Router Output) | Activated Tools in Sandbox | Mechanism & Description | Routing Rationale (Problem Solved) |
|---|---|---|---|
| Entity Disambiguation | Entity-disambiguation, network-exploration, event-history, and lexical full-text endpoints | Searches for Company and Person entities through fuzzy matching over names, locations, and purposes. It excludes orphan weak nodes through the NameHub hierarchy. | Supports initial entity identification while retaining the network and history endpoints required for the next analytical step. |
| Network Exploration | Entity-disambiguation, network-exploration, event-history, and lexical full-text endpoints | Executes a three-branch UNION ALL Cypher query to identify corporate structures, event subjects, and lateral connections across NameHub aliases. | Keeps entity search available because network traversal requires a resolved target entity. |
| Event-History Retrieval | Entity-disambiguation, network-exploration, event-history, and lexical full-text endpoints | Traverses an entity through its NameHub and collects chronologically ordered Event nodes associated with its legal aliases. | Retains entity search so that the agent can resolve the target UID before retrieving its timeline. |
| Macro-Level Analytics | Top-entity aggregation, event-conditioned counting, and read-only custom-query endpoints | Provides statistical summaries by ranking entities by selected metrics (e.g., risk rank or nominal capital) or counting nodes associated with specific sub-rubrics. | Removes entity-search and traversal endpoints so that aggregation is performed by predefined server-side queries rather than manual pagination. |
| Deep Text Fallback (Nested in non-analytic intents) | Lexical full-text search endpoint | Runs a broad lexical search utilizing a Neo4j Lucene full-text inverted index directly against raw legal text snippets. | Provides a fallback when structured entity matching fails for an uncommon string or an entity that is not represented as a resolved graph node. |
| Unrestricted Multi-Hop Exploration ( all ) | All 7 tools loaded simultaneously | Grants the agent access to the standard retrieval suite, analytics endpoints, and the read-only Cypher sandbox. | Supports ambiguous or multi-part questions that cannot be assigned reliably to a narrower intent category. |
| Metric | What it measures | Why it matters | Scoring prompt |
|---|---|---|---|
| Faithfulness | Whether the agent’s answer is fully grounded in the retrieved context, meaning that every factual claim in the answer can be traced back to the data the system actually retrieved from the database. | A fluent response may still contain claims that are absent from the retrieved evidence. A score of 1.0 indicates that all factual claims are supported by the context, whereas 0.0 indicates that the response is unsupported. This metric is used for the graph-seeded automated benchmark, where grounding to retrieved evidence is central to the evaluation. | “Given the context and the answer, compute a faithfulness score between 0.0 and 1.0. 1.0 means the answer is 100% derived from the context. 0.0 means the answer states completely fabricated facts. Output ONLY a float number.” |
| Correctness | Whether the agent’s answer is factually consistent with the verified reference answer, regardless of how the supporting context was retrieved. | Correctness is used for the manually curated benchmark, where each question has a verified reference answer. It measures factual agreement with that answer rather than only support from the retrieved context. | “Given the question, the expected answer, and the actual answer, compute a correctness score between 0.0 and 1.0. 1.0 means the actual answer is factually correct and captures the expected answer well. 0.0 means the actual answer is factually wrong or fails to answer the question correctly. Be tolerant to differences in wording, formatting, ordering, or level of detail, as long as the substance is correct. Output ONLY a float number.” |
| Answer Relevance | Whether the answer directly and completely addresses the user’s question, regardless of the retrieved context. | A factually supported response can still address the wrong question or obscure the requested information with irrelevant detail. Answer Relevance isolates whether the response directly addresses the question. | “Given the question and the answer, compute an answer relevance score between 0.0 and 1.0. 1.0 means it perfectly answers the question directly. 0.0 means it dodges the question entirely. Output ONLY a float number.” |
| Information Recall | The fraction of the expected answer’s core information that the agent’s response successfully captures. | Grounding, correctness, and relevance do not establish that an answer is complete. Information Recall compares the response with the reference answer and measures how much of its core information is present. It is therefore sensitive to incomplete retrieval and incomplete synthesis. | “Compare the Expected Answer to the Actual Answer. Score from 0.0 to 1.0 how much of the Expected Answer’s core information is present in the Actual Answer. Output ONLY a float number.” |
| Metric | What it measures | Why it matters | Scoring prompt / logic |
|---|---|---|---|
| Turn Success Rate | The fraction of turns in a conversation that are successfully answered according to the conversational scoring rule. | A system may answer the final question correctly after making errors in earlier turns. Turn Success Rate captures consistency across the full interaction and reflects the propagation of local errors through later turns. | For each turn, success is assigned deterministically after the judge scores the answer. A turn is marked successful if the normalized exact match equals 1.0, or if Correctness , or if Correctness and Answer Relevance . The final Turn Success Rate is the average of these binary turn outcomes over all turns in the conversation. |
| Context Carryover Accuracy | Whether the system correctly preserves the active entity, prior references, and relevant facts across follow-up turns that require memory of earlier dialogue state. | Conversational systems may answer isolated questions correctly but lose track of the active person, company, event, or previously retrieved fact. This metric evaluates continuity across follow-up questions involving counting, summarization, or co-reference resolution. | This metric is computed deterministically over the subset of turns tagged for context carryover, entity memory, or temporal memory. For these turns, the same binary success rule used in Turn Success Rate is applied. Context Carryover Accuracy is then defined as the average success over this memory-sensitive subset only. |
| Tool Transition Accuracy | Whether the GraphRAG Agent follows the expected structured tool sequence for a given conversational turn. | The graph-based architecture assigns different tools to different semantic intents. Tool Transition Accuracy evaluates whether the agent follows the expected structured retrieval path, thereby testing the state machine and sequential tool logic. | This metric is only computed for the agentic GraphRAG. For each turn that specifies an expected tool behavior, the execution trace is inspected to determine whether at least one of the expected tools was actually called. A hit receives a score of 1.0 and a miss receives 0.0. Tool Transition Accuracy is the average of these binary hits across all turns for which an expected tool behavior is defined. |
| Metric & Score | Question | Reference Answer | Agent Answer |
|---|---|---|---|
| Answer Relevance Score: 0.10 | What is the registered address of the company with anonymized UID CHE-000.000.000? | Street Name 19, 5300 City A | The registered address of the company Company A GmbH (UID: CHE-000.000.000) has changed to City A as of January 12, 2024. Previously, it was located in City B . Would you like to investigate anything else related to this company? |
| Information Recall Score: 0.55 | Within the NameHub ’Person X’, do the company links resolve to one or multiple distinct person profiles, and for each profile (by person_uid), which company is it associated with? | Multiple distinct profiles: two in total. weak_1 is associated with Company B Sàrl; weak_2 is associated with Foundation C. | The search for ”Person X” returned a weak node, which has been explored to reveal connections to other individuals. Here are the distinct person profiles associated with Person X, along with the companies they are linked to: • Person X: Company B Sàrl (Role: Associate) • Person X: Foundation C (Role: Board Member) Note: This entity is connected to a NameHub. You can open the NameHub to explore other records that might belong to the same underlying name cluster. Would you like to investigate further into any specific company or person? |
| Metric & Score | Question | Reference Answer | Agent Answer |
|---|---|---|---|
| Answer Relevance Score: 0.40 | Show all the companies with which Person Y is connected. | Person Y is connected with the following companies: Person Y, Company D Sàrl. | The entities shown below are connected to various records sharing this exact name, but the database cannot determine with 100% certainty that they belong to the exact same physical person. • Company D Sàrl (Location: City C, Connection: HAS_EVENT, Date: 2023-12-20) • Person Y (Location: City C, Connection: HAS_EVENT, Date: 2023-12-12) Would you like to investigate one of these companies further? |
| Information Recall Score: 0.00 | Have any notices of dissolution or bankruptcy been published for Company E? | Yes, a notice of dissolution for Company E was published on December 17, 2020. | The company ”Company E” located in City D has a history of bankruptcy. Here are the relevant details: 1. Bankruptcy Notice: The bankruptcy was officially opened by the Regional Court F, effective from December 10, 2020, at 14:02. 2. Previous Changes: On June 5, 2020, the company changed its name from ”Company G” to ”Company E” and relocated to Street Name 10, City D. 3. Initial Registration: The company was initially registered as ”Company G” on September 19, 2019, in City E. If you need more information or further investigation into related entities or events, please let me know! |
| Metric | Value |
|---|---|
| Labeled examples | 40 |
| Pearson correlation | 0.721 |
| Spearman correlation | 0.752 |
| Mean absolute error | 0.168 |
| Within-half-point agreement | 0.875 |
| Quadratic weighted Cohen’s | 0.702 |
| Support label | Count | Share |
|---|---|---|
| Supported | 14 | 0.233 |
| Partially supported | 46 | 0.767 |
| Not supported | 0 | 0.000 |
| Uncertain | 0 | 0.000 |
| Supported or partially supported | 60 | 1.000 |
| Architecture variant | Family | Controlled configuration |
| Dense Vector-RAG | Flat retrieval | Dense retrieval over event documents using the Chroma index. |
| Lexical Full-Text RAG | Flat retrieval | Neo4j full-text index retrieval at the selected setting. |
| Hybrid Dense+Lexical RAG | Flat retrieval | Reciprocal rank fusion of Chroma dense and Neo4j lexical candidates at . |
| GraphRAG w/o reflection | Graph ablation | Full graph tools and router retained, but bounded recovery is limited to one retrieval/planning attempt. |
| GraphRAG w/o router | Graph ablation | Reflection and graph tools retained, but the agent receives the full tool set for every query. |
| Structured-only GraphRAG | Graph ablation | Router and reflection retained, but LLM-extracted weak nodes are ignored during retrieval. |
| Architecture variant | N | Faith. | 95% CI | Ans. rel. | 95% CI | Info. recall | 95% CI |
|---|---|---|---|---|---|---|---|
| Dense Vector-RAG | 300 | 0.905 | [0.877, 0.931] | 0.081 | [0.055, 0.109] | 0.066 | [0.043, 0.090] |
| Lexical Full-Text RAG | 300 | 0.975 | [0.959, 0.988] | 0.047 | [0.026, 0.072] | 0.030 | [0.016, 0.046] |
| Hybrid Dense+Lexical RAG | 300 | 0.879 | [0.848, 0.908] | 0.127 | [0.094, 0.164] | 0.091 | [0.065, 0.119] |
| GraphRAG w/o reflection | 300 | 0.894 | [0.866, 0.923] | 0.425 | [0.373, 0.477] | 0.388 | [0.339, 0.438] |
| GraphRAG w/o router | 300 | 0.887 | [0.865, 0.908] | 0.742 | [0.704, 0.782] | 0.581 | [0.540, 0.625] |
| Structured-only GraphRAG | 300 | 0.865 | [0.838, 0.892] | 0.651 | [0.607, 0.696] | 0.534 | [0.486, 0.582] |
| Architecture variant | N | Corr. | 95% CI | Ans. rel. | 95% CI | Info. recall | 95% CI |
|---|---|---|---|---|---|---|---|
| Dense Vector-RAG | 60 | 0.143 | [0.075, 0.213] | 0.247 | [0.152, 0.346] | 0.118 | [0.064, 0.181] |
| Lexical Full-Text RAG | 60 | 0.237 | [0.147, 0.336] | 0.322 | [0.212, 0.437] | 0.261 | [0.158, 0.370] |
| Hybrid Dense+Lexical RAG | 60 | 0.262 | [0.176, 0.357] | 0.470 | [0.358, 0.581] | 0.237 | [0.157, 0.324] |
| GraphRAG w/o reflection | 60 | 0.552 | [0.428, 0.676] | 0.563 | [0.444, 0.680] | 0.560 | [0.435, 0.685] |
| GraphRAG w/o router | 60 | 0.708 | [0.598, 0.813] | 0.777 | [0.677, 0.868] | 0.703 | [0.592, 0.812] |
| Structured-only GraphRAG | 60 | 0.545 | [0.429, 0.659] | 0.733 | [0.629, 0.825] | 0.578 | [0.461, 0.695] |
| Architecture variant | N | Corr. | 95% CI | Ans. rel. | 95% CI | Turn succ. | 95% CI | Carryover | 95% CI |
|---|---|---|---|---|---|---|---|---|---|
| Dense Vector-RAG | 10 | 0.344 | N/A | 0.644 | N/A | 0.367 | [0.192, 0.558] | 0.433 | [0.250, 0.633] |
| Lexical Full-Text RAG | 10 | 0.342 | N/A | 0.493 | N/A | 0.308 | [0.175, 0.483] | 0.400 | [0.233, 0.567] |
| Hybrid Dense+Lexical RAG | 10 | 0.310 | N/A | 0.677 | N/A | 0.308 | [0.142, 0.508] | 0.350 | [0.183, 0.533] |
| GraphRAG w/o reflection | 10 | 0.458 | N/A | 0.598 | N/A | 0.450 | [0.292, 0.617] | 0.517 | [0.350, 0.683] |
| GraphRAG w/o router | 10 | 0.554 | N/A | 0.714 | N/A | 0.542 | [0.367, 0.708] | 0.517 | [0.350, 0.683] |
| Structured-only GraphRAG | 10 | 0.604 | N/A | 0.827 | N/A | 0.608 | [0.400, 0.808] | 0.650 | [0.467, 0.833] |
| Architecture variant | N | Evidence | 95% CI | Supported | 95% CI | Empty | 95% CI | Tool/error | 95% CI |
|---|---|---|---|---|---|---|---|---|---|
| Dense Vector-RAG | 60 | 0.283 | [0.167, 0.400] | 0.067 | [0.017, 0.133] | 0.000 | [0.000, 0.000] | 0.000 | [0.000, 0.000] |
| Lexical Full-Text RAG | 60 | 0.300 | [0.183, 0.417] | 0.133 | [0.050, 0.217] | 0.667 | [0.550, 0.783] | 0.000 | [0.000, 0.000] |
| Hybrid Dense+Lexical RAG | 60 | 0.533 | [0.417, 0.650] | 0.167 | [0.083, 0.267] | 0.000 | [0.000, 0.000] | 0.000 | [0.000, 0.000] |
| GraphRAG w/o reflection | 60 | 0.583 | [0.450, 0.717] | 0.550 | [0.433, 0.683] | 0.133 | [0.050, 0.217] | 0.150 | [0.067, 0.250] |
| GraphRAG w/o router | 60 | 0.850 | [0.750, 0.933] | 0.683 | [0.567, 0.800] | 0.000 | [0.000, 0.000] | 0.117 | [0.050, 0.200] |
| Structured-only GraphRAG | 60 | 0.783 | [0.667, 0.883] | 0.483 | [0.350, 0.617] | 0.000 | [0.000, 0.000] | 0.450 | [0.333, 0.567] |
| Result family | Primary recorded fields | Interpretation |
|---|---|---|
| Flat dense/lexical/hybrid runs | Retrieved-document counts, generated-answer text, explicit no-answer strings, API/judge errors | Captures empty retrieval, empty or explicit no-answer outputs, and execution errors; it does not represent graph-tool failures. |
| Graph-seeded graph-agent runs | Recorded trace logs, failed or empty tool-call indicators, recovery flags, generated answers, judge scores | Captures graph-tool retrieval failures or empty intermediate results; these may occur even when the final answer still receives a valid judge score. |
| Manually curated benchmark graph-agent runs | Recorded tool traces, answer text, judge scores, latency, and ablation-specific execution fields | Captures the same graph-tool reliability signals on the manually curated benchmark and keeps failed rows in aggregate quality means when scored. |
| Conversational runs | Turn-level answer text, trace fields, explicit no-answer strings, and graph-tool transition records | Captures turn-level unsuccessful outcomes; aggregate conversational metrics are computed separately at the conversation/turn level. |
| Architecture variant | Failed | Rate | Wilson 95% CI | Non-fail rel. | Non-fail recall |
|---|---|---|---|---|---|
| Dense Vector-RAG | 0/300 | 0.000 | [0.000, 0.013] | 0.081 | 0.066 |
| Lexical Full-Text RAG | 282/300 | 0.940 | [0.907, 0.962] | 0.791 | 0.494 |
| Hybrid Dense+Lexical RAG | 244/300 | 0.813 | [0.765, 0.853] | 0.673 | 0.484 |
| GraphRAG w/o reflection | 16/300 | 0.053 | [0.033, 0.085] | 0.448 | 0.410 |
| GraphRAG w/o router | 84/300 | 0.280 | [0.232, 0.333] | 0.821 | 0.723 |
| Structured-only GraphRAG | 99/300 | 0.330 | [0.279, 0.385] | 0.821 | 0.758 |
| Comparison | N | Diff. | 95% CI | A fail/F pass | A pass/F fail | Holm |
|---|---|---|---|---|---|---|
| Dense Vector-RAG minus Full | 300 | -0.073 | [-0.103, -0.047] | 0 | 22 | |
| Lexical Full-Text RAG minus Full | 300 | 0.867 | [0.823, 0.903] | 262 | 2 | |
| Hybrid Dense+Lexical RAG minus Full | 300 | 0.740 | [0.687, 0.790] | 223 | 1 | |
| GraphRAG without reflection minus Full | 300 | -0.020 | [-0.053, 0.013] | 9 | 15 | 0.307 |
| GraphRAG without intent routing minus Full | 300 | 0.207 | [0.160, 0.253] | 64 | 2 | |
| Structured-only GraphRAG minus Full | 300 | 0.257 | [0.207, 0.307] | 78 | 1 |
| Architecture variant | Empty retrieval | Empty answer | Retrieval exec. | Recovery flag | No-answer |
| Dense Vector-RAG | 0 | 0 | 0 | 0 | 0 |
| Lexical Full-Text RAG | 279 | 282 | 0 | 0 | 282 |
| Hybrid Dense+Lexical RAG | 0 | 244 | 0 | 0 | 244 |
| GraphRAG w/o reflection | 16 | 0 | 0 | 16 | 0 |
| GraphRAG w/o router | 84 | 0 | 0 | 84 | 0 |
| Structured-only GraphRAG | 98 | 0 | 1 | 99 | 3 |
| Architecture variant | Failed rows | Faith. failed | Faith. non-failed | Failed faith. 0.9 | Failed rel./recall 0.1 |
| Dense Vector-RAG | 0 | N/A | 0.905 | 0 | 0/0 |
| Lexical Full-Text RAG | 282 | 0.995 | 0.648 | 280 | 282/282 |
| Hybrid Dense+Lexical RAG | 244 | 0.882 | 0.863 | 200 | 243/243 |
| GraphRAG w/o reflection | 16 | 0.131 | 0.937 | 2 | 16/16 |
| GraphRAG w/o router | 84 | 0.841 | 0.904 | 55 | 13/35 |
| Structured-only GraphRAG | 99 | 0.726 | 0.934 | 53 | 46/81 |
| Architecture variant | Grouping | Value | N | Failed | Rate | Non-fail recall |
|---|---|---|---|---|---|---|
| Dense Vector-RAG | difficulty | Level 1 | 100 | 0 | 0.000 | 0.078 |
| Dense Vector-RAG | difficulty | Level 2 | 100 | 0 | 0.000 | 0.070 |
| Dense Vector-RAG | difficulty | Level 3 | 100 | 0 | 0.000 | 0.049 |
| Dense Vector-RAG | task type | Direct Entity Data | 100 | 0 | 0.000 | 0.078 |
| Dense Vector-RAG | task type | NameHub Entity Resolution | 100 | 0 | 0.000 | 0.070 |
| Dense Vector-RAG | task type | Temporal History | 100 | 0 | 0.000 | 0.049 |
| Bench. | Comp. | Metric | N | Delta [95% CI] | |
| GS | FA–DV | Faithfulness | 300 | 0.704 | |
| GS | FA–DV | Answer Relevance | 300 | ||
| GS | FA–DV | Information Recall | 300 | ||
| GS | FA–LF | Faithfulness | 300 | ||
| GS | FA–LF | Answer Relevance | 300 | ||
| GS | FA–LF | Information Recall | 300 |
| Faith. | Corr. | Rel. | Recall | Mean s | Median s | P95 s | Avg. Docs | |
|---|---|---|---|---|---|---|---|---|
| 5 | 0.858 | 0.082 | 0.142 | 0.064 | 16.97 | 16.23 | 23.26 | 4.65 |
| 10 | 0.875 | 0.103 | 0.197 | 0.082 | 16.61 | 16.09 | 24.43 | 9.25 |
| 20 | 0.887 | 0.112 | 0.201 | 0.086 | 16.86 | 16.44 | 24.17 | 18.50 |
| Faith. | Corr. | Rel. | Recall | Mean s | Median s | P95 s | Avg. Docs | |
|---|---|---|---|---|---|---|---|---|
| 5 | 0.974 | 0.134 | 0.195 | 0.117 | 13.57 | 13.24 | 19.04 | 1.67 |
| 10 | 0.992 | 0.215 | 0.312 | 0.221 | 16.53 | 15.67 | 28.37 | 2.98 |
| 20 | 0.977 | 0.237 | 0.322 | 0.261 | 14.29 | 13.92 | 21.05 | 4.62 |
| Faith. | Corr. | Rel. | Recall | Mean s | Median s | P95 s | Avg. Docs | |
|---|---|---|---|---|---|---|---|---|
| 5 | 0.909 | 0.195 | 0.310 | 0.171 | 20.13 | 17.73 | 34.26 | 5.00 |
| 10 | 0.963 | 0.203 | 0.404 | 0.169 | 16.78 | 15.67 | 24.13 | 10.00 |
| 20 | 0.958 | 0.262 | 0.470 | 0.237 | 17.50 | 16.62 | 24.78 | 20.00 |
| Exact Hit | Token Recall | Token Hit@50% | Nonempty | |
|---|---|---|---|---|
| 5 | 0.000 | 0.259 | 0.217 | 1.000 |
| 10 | 0.000 | 0.279 | 0.233 | 1.000 |
| 20 | 0.000 | 0.300 | 0.233 | 1.000 |
| Metric | Operational definition | Value |
|---|---|---|
| Evidence-bearing retrieval success | Evidence in retrieved context | 0.9500 |
| Correct-answer supported trajectory | Correctness and evidence-bearing | 0.8000 |
| Non-evidence retrieval rate | Returned data, but no answer-bearing evidence | 0.0500 |
| Tool error / empty-result rate | Empty/error tool calls over all tool calls | 0.0132 |
| Reflection recovery rate | Later evidence after weak first retrieval | 0.6000 |
| Duplicate tool-call rate | Repeated same tool and arguments | 0.0000 |
| Extracted item | Precision | Recall | F1 | Matching criterion |
|---|---|---|---|---|
| Persons | 0.989 | 0.956 | 0.972 | Fuzzy name match |
| Companies | 0.933 | 0.651 | 0.767 | Fuzzy name match, excluding structured subject entities |
| Roles | 0.723 | 0.611 | 0.662 | Coarse-compatible role match |
| Entity-role relations | 0.701 | 0.482 | 0.571 | Coarse-compatible entity-role match |
| Language | Persons | Companies | Roles | Relations | |
|---|---|---|---|---|---|
| French | 36 | 1.000/1.000/1.000 | 1.000/0.846/0.917 | 0.800/0.696/0.744 | 0.765/0.542/0.634 |
| German | 36 | 1.000/0.929/0.963 | 0.714/0.556/0.625 | 0.829/0.723/0.773 | 0.818/0.540/0.651 |
| Italian | 33 | 0.966/0.966/0.966 | 1.000/0.538/0.700 | 0.658/0.532/0.588 | 0.571/0.400/0.471 |
| For 15 notices, the archived validation records did not provide a reliable language label, and the text-based heuristic could not assign one confidently. These notices are included in the aggregate 120-notice results but excluded from this language-specific comparison. | |||||
| Metric | Value |
|---|---|
| Company nodes evaluated | 1,039,198 |
| Unique official UIDs | 806,520 |
| NameHubs evaluated | 908,203 |
| True-positive pairs | 114,556 |
| False-positive pairs | 41,837 |
| False-negative pairs | 157,006 |
| Dataset | Metric | Graph Mean [95% CI] | 95% CI | ||
|---|---|---|---|---|---|
| GS | Faithfulness | 0.898 [0.874, 0.920] | -0.007 | [-0.044, 0.029] | 0.693 |
| GS | Answer Relevance | 0.689 [0.646, 0.730] | 0.609 | [0.557, 0.658] | |
| GS | Information Recall | 0.574 [0.528, 0.618] | 0.508 | [0.456, 0.561] | |
| MC | Correctness | 0.828 [0.735, 0.913] | 0.685 | [0.565, 0.797] | |
| MC | Answer Relevance | 0.887 [0.822, 0.941] | 0.641 | [0.517, 0.756] | |
| MC | Information Recall | 0.838 [0.746, 0.919] | 0.719 | [0.611, 0.819] |
| Statistic | Value |
|---|---|
| Conversations | 10 |
| Total turns | 36 |
| Mean turns | 3.60 |
| Minimum turns | 3 |
| Maximum turns | 4 |
| Std. turns | 0.49 |
| Dataset | System | Mean s | Median s | P95 s | Mean Tools | Max Tools | |
|---|---|---|---|---|---|---|---|
| GS | Full Agentic GraphRAG | 300 | 12.53 | 11.36 | 26.93 | 1.67 | 4 |
| GS | Dense Vector-RAG | 300 | 32.45 | 32.35 | 37.66 | – | – |
| MC | Full Agentic GraphRAG | 60 | 10.27 | 9.90 | 20.11 | 1.27 | 3 |
| MC | Dense Vector-RAG | 60 | 15.62 | 15.39 | 27.92 | – | – |
| CONV | Full Agentic GraphRAG | 10 | 30.60 | 28.37 | 61.63 | 3.70 | 7 |
| CONV | Dense Vector-RAG | 10 | 37.06 | 35.06 | 58.64 | – | – |
| Category | Implementation identifier | Reader-facing terminology |
|---|---|---|
| Agent endpoint | search_companies | Entity-disambiguation endpoint |
| Agent endpoint | explore_network | Network-exploration endpoint |
| Agent endpoint | get_node_history | Event-history endpoint |
| Agent endpoint | global_text_search | Lexical full-text search endpoint |
| Agent endpoint | get_top_entities | Top-entity aggregation endpoint |
| Agent endpoint | count_entities_by_event | Event-conditioned counting endpoint |