From Retrieval to Customer Context: Evaluating Frontier-Model Systems for Voice-of-Customer Analysis
Organizations: Enterpret
Abstract
Organizations increasingly use frontier language models to analyze customer feedback, but answer quality also depends on how that feedback is organized and made available. We define a \emph{customer context graph} as a unified model of customer and business context. Typed relationships connect customer objects (feedback, conversations, users, and accounts), operational objects (tickets, support agents, opportunities, and competitors), and analytical or action objects (taxonomy concepts, evidence, insights, work items, and outcomes). This lets an agent investigate not only what customers say, but why, who is affected, what action followed, who owns it, and whether it was resolved. For this experiment, the graph is populated from public Cursor feedback; the same architecture can support any type of feedback source. We compare Agentic RAG, a Deep Research Agent, and a Customer Context Graph-backed Agent on the same 9,432 public Cursor feedback records using 30 realistic product, incident, comparison, and metadata questions. Without exhaustive ground truth, we jointly score responses on answer quality (coverage and organization), analytical depth (specificity and decomposition), and evidence quality (citation support and traceability), using a comparative rubric calibrated on 28 of the 30 questions. We sample cited records against their claims and weight the three dimensions equally. Under this aligned rubric, the Customer Context Graph-backed Agent scores 0.961 overall, versus 0.710 for the Deep Research Agent and 0.651 for Agentic RAG, and leads the Deep Research Agent on 27 of 30 paired questions (sign-test p < 10^(-5); strictly best on 26 of 30). Its largest advantage is analytical depth (0.967 versus 0.642), reflecting more specific, hierarchically developed findings with quantified themes and traceable evidence...
Figures & tables
| Held constant | Frozen 9,432-record corpus + SHA-256; identical query text and as-of date; base analyst model, both judges, retrieval reranker; answer contract; retry policy; alias-shuffled evaluator presentation. |
| Allowed to differ (the treatment) | Whether the system reconstructs product concepts per query or reads pre-computed summaries, taxonomy predictions, metadata, user information, and graph relationships stored at ingestion; the orchestration budget the analyst model is dispatched into. |
| Can therefore claim | The complete-system effect of persistent customer context on this corpus, question set, and calibrated rubric; not the isolated causal effect of any component in isolation. |
| Statistic | Value | Coverage |
|---|---|---|
| Normalized feedback records | 9,432 | 100% |
| Collection window | 2026-04-16–2026-07-15 UTC | |
| Public feedback sources | 3 | 100% |
| Records with rating | 171 | 1.81% |
| Records with normalized version | 5 | 0.05% |
| Records with public country | 118 | 1.25% |
| Depth | Required reasoning | |
|---|---|---|
| 1 | Exact lookup, metadata coverage, capability checks | 11 |
| 2 | Focused semantic classification or evidence synthesis | 10 |
| 3 | Comparisons across source, time, or cohort | 3 |
| 4 | Multi-stage VoC, strategy, incident, and workflow decisions | 6 |
| Total | 30 |
| System | Overall (95% CI) | Answer quality | Analytical depth | Evidence quality | Queries |
|---|---|---|---|---|---|
| Agentic RAG | 0.651 [0.613, 0.692] | 0.625 | 0.642 | 0.686 | 30/30 |
| Deep Research Agent | 0.710 [0.668, 0.752] | 0.683 | 0.642 | 0.806 | 30/30 |
| Context Graph Agent | 0.961 [0.936, 0.983] | 0.958 | 0.967 | 0.958 | 30/30 |
| Comparison | Metric | Wins | Ties | Losses | (95% CI) | Sign-test |
|---|---|---|---|---|---|---|
| Context graph Deep | Overall | 27 | 1 | 2 | +0.251 [+0.193, +0.305] | |
| Answer quality | 25 | 1 | 4 | +0.275 | ||
| Analytical depth | 28 | 0 | 2 | +0.325 | ||
| Evidence quality | 27 | 2 | 1 | +0.153 | ||
| Context graph Agentic | Overall | 29 | 0 | 1 | +0.310 [+0.263, +0.354] | |
| Answer quality | 28 | 1 | 1 | +0.333 |
| System | Cited | Member-audited | Traceable aggregate | Unresolved | Records sampled |
|---|---|---|---|---|---|
| Agentic RAG | 712 | 708 | 0 | 4 | 708 |
| Deep Research Agent | 594 | 594 | 0 | 0 | 594 |
| Context Graph Agent | 913 | 127 | 786 | 0 | 127 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Abridged evaluated response | Overall |
|---|---|---|
| Agentic RAG | Runs a model-coded exclusive classification over 8,890 of 9,432 records (94.25% coverage; 2,728 assigned to “other/no clear pain point”) and ranks seven pain-point categories: billing, pricing, usage limits, credits, overages, and refunds (2,195); model choice, Auto-mode opacity, and context/routing (1,461); agent/code quality and unsafe edits (875); trust, privacy, and ownership concerns (748); product instability, regressions, and IDE performance (629); mobile/iOS parity gaps (155); and support/account escalation (99). Each category is followed by cited examples and an analyst interpretation. | 0.556 |
| Deep Research Agent | Runs an exhaustive semantic classification over all 9,432 records into ten exclusive primary pain-point categories and reports both counts and corpus share: model quality and Auto-mode control (892; 9.5%); token burn and unclear usage attribution (628; 6.7%); usage caps and quota lockout (370; 3.9%); billing surprises and commercial trust (353; 3.7%); long-session context loss (232; 2.5%); update regressions (207; 2.2%); UI churn (156; 1.7%); agent out-of-scope changes (119; 1.3%); support and refund friction (116; 1.2%); and crashes or severe local instability (79; 0.8%). It also flags 2,133 records classified as positive value or mitigation as a counter-signal. | 0.776 |
| Context Graph Agent | Ranks recurring themes over 5,035 customer conversations (Reddit 4,868 at 96.7%; App Store 114; Trustpilot 53) organized into seven clusters: quota and usage-limit friction (180), throughput rate-limit disruption (162), degraded behavior when rate limits are hit (95), Auto routing mismatch (92), Composer/model quality regressions (86), poor AI output quality (82), and a service-availability/agent-safety cluster whose top themes are pre-execution safety controls (77) and service availability issues (77). Each main count links to its evidence set; leading themes decompose into cited subtheme drivers and are illustrated with verbatim customer quotes. | 1.000 |
| System | Abridged evaluated response | Overall |
|---|---|---|
| Agentic RAG | Compares 30-day calendar windows (preceding 2026-05-17 to 2026-06-15 with 557 records; recent 2026-06-16 to 2026-07-15 with 8,652 records) and finds that a keyword reliability proxy falls in share from 17.2% to 9.8% even as absolute counts rise. The only rising sub-signal is a regression/update-language proxy that grows from 1 to 37 records, and 113 App Store reviews appear only in the recent window (no preceding baseline). It labels the broad “reliability worsening” claim as low confidence. | 0.556 |
| Deep Research Agent | Uses closely matched rolling 14-day windows (3,720 versus 3,717 records). Its combined reliability proxy rises from 301 to 345 (+14.6%), while native-tool, harness, MCP, and file-operation language rises from 79 to 132 (+67.1%). It localizes a GLM 5.2 tool-call/markup cluster to 2026-07-08 and 2026-07-09 (21 of 25 recent GLM matches), distinguishes increasing, flat, and decreasing signals across category proxies, and assigns medium confidence. | 0.930 |
| Context Graph Agent | Reports a concrete recent-period reliability signal across 9,408 conversations (May 31 to Jul 15) with 86 reliability-adjacent themes and five failure clusters (rate-limit and quota reliability, agent execution stalls, core product reliability, billing/metering errors, client-side crashes and freezes), each with citation-linked recent and preceding counts. It declines to claim an increase because the preceding period (Apr 16 to May 30) contains only 24 Trustpilot-only conversations while the recent period is 98.4% Reddit; the only apples-to-apples slice is Trustpilot’s move. | 0.786 |
| System | Abridged evaluated response | Overall |
|---|---|---|
| Agentic RAG | States the denominator (63 of 113 App Store reviews; 54 five-star and 9 four-star; 2026-06-29 to 2026-07-13) and organizes praise into five specific, separately developed themes—mobility (“I can work anywhere now”), cloud-agent control as the primary use case, concrete productivity gains from fewer idle agents, focused mobile UX, and continuity with desktop workflows. Each theme carries several directly quoted reviews with record citations and closes with an explicit analyst interpretation distinguishing what is being praised (a control plane for agents) from what is not (mobile code editing). | 0.887 |
| Deep Research Agent | Establishes the same denominator with an auditable exact query and rating-coverage checks (63 of 113; 55.8%). Reports five praise themes but scores them as retrieved-evidence counts over a 39-record retrieval set (mobile agent work 18/39; interface polish 8/39; agent visibility 6/39; pre-existing Cursor affinity 7/39; mixed praise-plus-request 10/39), and states plainly that no exhaustive classification was scoped to the 63 high-rated reviews, so these are not prevalence estimates. The disclosure is exemplary but the resulting themes stay closer to retrieval bookkeeping than to decision-ready analysis. | 0.608 |
| Context Graph Agent | Reports the same 63-conversation denominator with a drill-down citation, and is the only system to flag that App Store coverage begins only in late June 2026 (an ingestion gap rather than an absence of reviews). It returns five citation-linked praise themes with conversation counts—mobile agent companion (11), general product delight (8), usability and design (5), value of off-device agent execution (5), speed (2)—plus explicit mixed-intent and non-additive-count caveats. It earns perfect evidence quality (1.000) but tier-two answer quality and analytical depth (0.750 each): its second-largest cluster is the catch-all “general product delight — undifferentiated love for Cursor overall,” precisely the kind of umbrella label the rubric’s one-problem test penalizes, and its leading specific theme covers only 11 of 63 reviews. | 0.833 |
| Depth | Agentic | Deep Res. | Context graph | |
|---|---|---|---|---|
| 1 | 11 | 0.682 | 0.692 | 0.955 |
| 2 | 10 | 0.672 | 0.679 | 0.967 |
| 3 | 3 | 0.555 | 0.834 | 0.892 |
| 4 | 6 | 0.606 | 0.733 | 0.996 |
| Analysis family | Agentic RAG | Deep Research | Context graph | |
|---|---|---|---|---|
| Executive VoC | 5 | 0.554 | 0.800 | 0.974 |
| Product area | 5 | 0.616 | 0.764 | 0.951 |
| Reliability and incidents | 5 | 0.632 | 0.733 | 0.918 |
| Pricing and packaging | 5 | 0.700 | 0.633 | 0.996 |
| Source and metadata | 5 | 0.767 | 0.637 | 0.928 |
| Competition and retention | 4 | 0.635 | 0.693 | 1.000 |
| Dimension | What the evaluator checks |
|---|---|
| Answer quality | Explicit question coverage, organization, presentation, specificity, useful quantification, and suitability for the intended reader. |
| Analytical depth | Diversity and granularity of findings; coherent themes, drivers, subthemes, contexts, counterevidence, implications, and synthesis. Broad umbrella themes are disfavored when they merge distinct problems. |
| Evidence quality | Relevance of deterministically sampled cited records plus comparative evidence strength: claim-level support, denominators, count semantics, aggregate drill-down links, and representative record citations. |
| Overall | Unweighted mean of the three dimensions for each query, followed by macro-averaging across the 30 questions. |
| Role | Model | Provider | Sampling | Budget |
| Agentic RAG | ||||
| Analyst (planning, tool use, synthesis) | gpt-5.5 | OpenAI | Provider default | 50 turns |
| Retrieval reranker | rerank-2.5 | Voyage | Provider default | — |
| Deep Research Agent | ||||
| Triage router | gpt-5.5 | OpenAI | Provider default | 50 turns |
| Planner / replanner | gpt-5.5 | OpenAI | Provider default | 50 turns |
| Component | Required content |
|---|---|
| Run | Candidate, model, prompt/config hashes, corpus hash, timestamps, status |
| Scope | Query ID and requested output context |
| Answer | Final rendered text |
| Citation ledger | Citation ID, nearby claim, cited record or aggregate members |
| Usage | Model calls, tokens, tools, cost fields, failures |
| Family | Family | ||
|---|---|---|---|
| Executive VoC | 5 | Product area | 5 |
| Reliability and incidents | 5 | Pricing and packaging | 5 |
| Source and metadata | 5 | Competition and retention | 4 |
| Cross-functional action | 1 |