cs.CRSep 30, 2026

RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage

Authors: Zeming Liu, Qibai Chen, Jingtao Zhang, Hang Lyu

Organizations: Brown University, United States · Independent Researcher, United States · Georgia Institute of Technology, United States

Abstract

Retrieval-augmented generation (RAG) systems need inexpensive ways to route generated answers: accept low-risk outputs, review uncertain ones, and reserve strong verifiers for the expensive tail. We present RAGScope, a leakage-controlled protocol for evaluating local evidence gates that use only the task input, retrieved context, and answer text. The protocol combines context-grouped splits, fold-scoped preprocessing, group bootstrap intervals, deployment operating points, end-to-end runtime, and explicit source-shift stress tests. On three RAGTruth tasks, the enhanced gate RAGScope-E reaches 0.798 AUROC and 0.660 average precision (AP) in pooled grouped cross-validation. Its pooled AP exceeds ROUGE-L by 0.034 with a 95% context-group interval of [0.002, 0.064], although the AUROC gain is not significant and ROUGE-L remains stronger on data-to-text. At a top-10% review budget, RAGScope-E attains 0.748 precision; accepting the lowest-risk 50% yields 0.141 residual unfaithfulness. RAGScope-E runs in 6.22 ms/example on CPU, versus 145.75 and 223.07 ms/example for the tested DeBERTa-NLI and HHEM settings. A 14,900-example HaluBench stress test exposes the deployment boundary: an in-domain calibrated gate reaches 0.879 AUROC, but leave-source-out calibration averages only 0.466. Target-only calibration recovers to 0.675 AUROC with 100 labels per source and 0.685 with 200. Cheap evidence gates are therefore useful routing components, but learned calibration must be validated and adapted within the target domain.

Figures & tables

Explore similar work

CardsList
  1. SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation

    May 5, 2026Jingxi Qiu, Zeyu Han, Cheng HuangHievi-RagRetrieval-Augmented Generation Pipelines

  2. AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation

    Oct 1, 2026Giulio Zeloni, Enrico Lo Conte, Salvatore Rionero +3Agentic EvaluationsLarge Language Model Judges

  3. Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG

    Aug 3, 2026Daeyoung Roh, Donghee HanHievi-Rag