Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit
Authors: Dmitrij Żatuchin
Organizations: Department of Information Technologies, Estonian Entrepreneurship University of Applied Sciences (EUAS), Tallinn, Estonia · Rankfor.AI, Tallinn, Estonia
Repeated-query audits must distinguish recovery of a collected set from completeness of possible outputs. We apply sample-based rarefaction to 4,500 responses from 50 buying questions, six configurations and 15 calls per cell. Historical-dictionary median ten-call recovery of the observed 15-call set ranges from 92.6% to 95.2%; re-adjudicating all 45,683 candidate strings changes this range to 89.5%-94.7%. Two blinded Gemini 3.1 Pro annotation roles assessed 600 complete answers, yielding micro F1 of 0.908 for canonical-name agreement and 0.975 for span-overlap agreement. This is AI-based evidence, without a human reference study. A separate matched roster analysis of 3,750 records per wave gives median single-call recovery of the observed five-call set of 80.0%-92.5% in February and 90.0%-100.0% in September, with question-subset dependence. Source accumulation also changes when API-returned hosts are restricted to those referenced by answer citation markers. These findings show that recovery percentages depend on extraction, question selection and the finite reference collection. They support explicit measurement definitions and sensitivity analyses, without establishing exhaustive repertoires, causal retrieval effects or a universal stopping rule.
Figures & tables
Engine
Sobs
Q1>0
A(1)/A(5)
A(10)/A(15)
Claude Sonnet 5
15
45/50
0.708
0.930
Gemini 3.7 Flash
16.5
45/50
0.724
0.943
GPT-5.6-luna
16
46/50
0.647
0.926
Grok 4.5
19
43/50
0.672
0.943
Mistral Large
31
46/50
0.622
0.943
Sonar, search enabled
8
32/50
0.768
0.952
Table 1: Historical dictionary, September 4 experiment: 50 cells per engine and 15 runs per cell. All ratios and set sizes are medians over cells. Q1>0 counts cells with at least one singleton.
Figure 1: Expected organization discovery, conditional on the stored candidate-based extraction. Left: median distinct counts. Right: median recovery relative to each cell’s observed 15-run set. Curves describe the collected panel; they do not identify the fraction of the entire possible output repertoire.
A(10)/A(15)
Configuration
H
A
F
GPT-5.6-luna
0.926
0.913
0.910
Claude Sonnet 5
0.930
0.921
0.920
Gemini 3.7 Flash
0.943
0.934
0.936
Grok 4.5
0.943
0.926
0.923
Mistral Large
0.943
0.897
0.895
Table 2: Sensitivity on identical September 4 answers. H: historical dictionary; A: add-only restoration; F: full re-adjudication. All entries are medians over 50 cells per configuration.
Figure 2: Dictionary sensitivity on the same responses. The add-only arm preserves baseline errors; full re-adjudication also changes classifier context and organization mapping. Neither revised arm is an independently validated gold standard.
Strict canonical names
Literal span overlap
Dictionary
P
R
F1
P
R
F1
Historical
0.848
0.865
0.857
0.918
0.936
0.927
Add-only
0.847
0.877
0.861
0.915
0.947
0.931
Full re-adjudication
0.844
0.883
0.863
0.916
0.959
0.937
Table 3: Agreement with the AI reference on the same 600 answers. P: precision; R: recall; F1: harmonic mean. These are conditional model-reference scores, not human-ground-truth accuracy.
Configuration
Complete
Defined
February
September
Change
GPT-5.2
207
203
0.900
0.900
0.0
Gemini 3 Flash preview
250
237
0.925
0.900
0.0
Sonar Pro
250
221
0.800
1.000
+5.7
Table 4: Paired non-empty five-call cells, own-industry roster. Defined pairs have non-empty unions in both waves. Recovery medians use these same pairs. Changes are median within-pair differences, in percentage points.
Figure 3: Median finite-sample recovery of the observed five-call own-industry roster set, over the same defined pairs in each wave. A flat median does not mean every question has a flat curve.
Figure 4: Source-host accumulation. Left: four earlier 24-run cells using their recorded source sets. Right: the September Sonar panel, distinguishing all returned source metadata from hosts referenced by numeric markers in the answer. Each line represents the indicated operational definition.