cs.IRSep 4, 2026

Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit

Authors: Dmitrij Żatuchin

Organizations: Department of Information Technologies, Estonian Entrepreneurship University of Applied Sciences (EUAS), Tallinn, Estonia · Rankfor.AI, Tallinn, Estonia

Abstract

Repeated-query audits must distinguish recovery of a collected set from completeness of possible outputs. We apply sample-based rarefaction to 4,500 responses from 50 buying questions, six configurations and 15 calls per cell. Historical-dictionary median ten-call recovery of the observed 15-call set ranges from 92.6% to 95.2%; re-adjudicating all 45,683 candidate strings changes this range to 89.5%-94.7%. Two blinded Gemini 3.1 Pro annotation roles assessed 600 complete answers, yielding micro F1 of 0.908 for canonical-name agreement and 0.975 for span-overlap agreement. This is AI-based evidence, without a human reference study. A separate matched roster analysis of 3,750 records per wave gives median single-call recovery of the observed five-call set of 80.0%-92.5% in February and 90.0%-100.0% in September, with question-subset dependence. Source accumulation also changes when API-returned hosts are restricted to those referenced by answer citation markers. These findings show that recovery percentages depend on extraction, question selection and the finite reference collection. They support explicit measurement definitions and sensitivity analyses, without establishing exhaustive repertoires, causal retrieval effects or a universal stopping rule.

Figures & tables

Explore similar work

CardsList
  1. Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines

    Sep 21, 2026Aviral Joshi, Hanoz Bhathena, Max Nelson +1Retrieval-Augmented Generation PipelinesModel Auditing

  2. ProvenAI: Provenance-Native Traces of Evidence in Generated Answers

    Jun 24, 2026Mohammad Faizan, Dalal AlharthiCitationsTraceability