RAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk Constraints
Organizations: Independent open-source research project
Abstract
Retrieval-augmented generation systems are extensively instrumented with metrics, benchmarks, traces, and automated judges, but these tools do not decide whether a proposed policy change is safe to release. We present RAGWarrant, an open-source promotion-control framework that treats deployment as a constrained evidence decision rather than a leaderboard choice. RAGWarrant normalizes evaluator outputs and operational telemetry, applies predeclared quality and hard-risk gates, assigns evidence-class claim ceilings, preserves negative outcomes, and emits auditable PROMOTE, BLOCK, REJECT, or INCONCLUSIVE decisions. We evaluate the framework across T2-RAGBench, MultiHop-RAG, CRAG, HotpotQA, synthetic reproduction, and bounded local generative experiments. On HotpotQA, operational savings were blocked because answer quality fell beyond the declared margin. A bounded CRAG study selected a lower-cost quality-tied policy, but related generative gains were unstable and a held-out guardrail failed closed. We claim an auditable promotion-control abstraction, not optimizer superiority, human validation, or production readiness. The tagged artifact reproduces from a fresh clone, runs as a hardened Docker job, accepts external evaluator exports, and verifies artifact integrity.
Figures & tables
| System family | Primary emphasis | Typical output | RAGWarrant relationship |
|---|---|---|---|
| RAGAS / ARES / RAGChecker | Quality and diagnostic evaluation | Metric scores and component diagnoses | Consumes declared metrics as governance evidence |
| RAGBench / CRAG / HotpotQA | Benchmark data and labels | Examples, references, supporting facts, scores | Uses bounded evidence sources with dataset-specific claim ceilings |
| DSPy and optimizers | Search over programs or policies | Candidate maximizing an objective | Treats optimizer output as a candidate, not automatic promotion |
| LangSmith / Phoenix | Experiments, traces, observability | Runs, evaluator scores, operational telemetry | Normalizes exports and applies release gates |
| NIST AI RMF / ISO 42001 | Lifecycle and management governance | Practices, outcomes, management requirements | Implements a narrow technical promotion control |
| RAGWarrant | Evidence-preserving promotion control | PROMOTE, BLOCK, REJECT, or INCONCLUSIVE plus audit bundle | Coordinates evidence, gates, claim ceilings, and integrity |
| Decision | Meaning | Typical trigger | Consequence |
|---|---|---|---|
| PROMOTE | Evidence supports replacing the baseline within declared scope. | Quality and operational/risk objectives pass; no hard disqualifier. | Candidate may advance with recorded boundaries. |
| BLOCK | Promotion cannot be evaluated safely. | No usable signal, leakage, missing provenance, security or hygiene failure. | No release claim; blocker and remediation are recorded. |
| REJECT | Evidence favors retaining the baseline. | Confirmed quality loss, negative held-out result, or protected regression. | Candidate is not promoted; negative evidence remains append-only. |
| INCONCLUSIVE | Available evidence does not resolve the decision. | Interval crosses threshold, evidence is mixed, or sample is underpowered. | No promotion; additional evidence may be collected. |
| Dataset / path | Role | Evidence characteristics | Primary limitation |
|---|---|---|---|
| T2-RAGBench | Public end-to-end development | 1,142-query development corpus; behaviorally variable policies | Development evidence; deterministic extractive generation in key runs |
| MultiHop-RAG | Public confirmatory corpus | 331-query sealed confirmatory test; full corpus-backed path | Governance matched quality-only |
| RAGBench HotpotQA | Context-retrieval enablement | Policy-dependent context assembly | Not full source-corpus retrieval in packaged evidence |
| CRAG web documents | Full corpus-backed retrieval | 2,706 rows; 9,848 web documents; 571 confirmatory rows | Noncommercial restriction; governance matched quality-only |
| CRAG mock API | Tool routing and generative validation | Live and frozen paths; calls, cost, latency, evaluator mapping | Strongest positive result bounded; generative results unstable |
| HotpotQA | Alternate corpus with answer labels | 1,000 local examples; 249 confirmatory behavioral rows | Operational gain accompanied by quality loss |
| Evaluation | Scope / N | Primary outcome | Interpretation |
|---|---|---|---|
| MultiHop-RAG confirmatory | 331 queries | Governance delta 0; noninferior, not superior | Valid confirmatory execution; no governance advantage |
| CRAG web documents | 571 rows | Governed = quality-only = top_k_high | Corpus-backed execution; no superiority |
| CRAG mock API | 571 rows | Utility +0.00100254; 571/0/0 wins/ties/losses | Strongest bounded source/retrieval signal; largely cost/latency driven |
| Frozen behaviorally distinct CRAG | 571 paired observations | Quality -0.00515371; cost -2.36839; calls -1.79860 | Lower operating burden within declared proxy-quality margin; derived evidence |
| HotpotQA behavioral | 249 confirmatory rows | Quality -0.0408367; F1 -0.0843373; cost -2.33014; calls -1.98394 | Operational gain with quality loss; blocked |
| CRAG generative primary | 12 examples; 132 generations | Quality +0.0166534; cost -3.77900; latency -5,971.85 ms | Positive bounded slice; repeats/models did not reproduce |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Evidence class | What it establishes | Permitted claim example | Not permitted |
|---|---|---|---|
| Fixture / smoke | Code path, schema, and refusal behavior execute | Harness runs and artifacts validate. | Benchmark or quality superiority |
| Development | Candidate behavior on nonsealed data | Promising development signal. | Confirmatory or production claim |
| Public confirmatory | Held-out result under frozen configuration | Bounded external signal on this corpus. | Universal generalization |
| Frozen-observation derived | Ablation and counterfactual analysis | Derived operational comparison. | Independent replication |
| Generative local | Pinned generator and generated-answer scoring | Bounded local generative validation. | Human or platform validation |
| Human evaluation | Completed blinded annotations | Human-adjudicated result. | Production outcome without operations |
| Field group | Representative fields | Purpose |
|---|---|---|
| Identity | run_id, suite, timestamp_utc, schema_version | Trace the decision to code, config, and evidence. |
| Decision | decision, result_class, decision_reason | Separate runtime status from governance outcome. |
| Selection | selected_policy, baseline_policy | Record what was compared and what, if anything, advances. |
| Effects | quality_delta, cost_delta, latency_delta, evidence_support_delta | Expose the empirical basis. |
| Risk | risk_flags, claim_boundaries, validator_status | Prevent operational gains from overriding disqualifiers. |
| Artifacts | artifact_uris, manifest hashes | Support independent review and tamper detection. |
| Claim | Status at v0.1.1-rc1 | Evidence boundary |
|---|---|---|
| Open-source governance engine | Supported | Fresh clone, CLI, public mini, schemas, tests, Docker job |
| CRAG source/retrieval governance signal | Supported with boundaries | Mock-API result is cost/latency driven and uses a restricted dataset |
| Selector governance blocks unsafe choices | Supported as stress-test evidence | Packaged cases; not a population estimate |
| RAG Compass superiority | Unsupported | Ranks below alternatives in multiple runs |
| Stable generative cost/latency superiority | Unsupported | Positive primary slice did not replicate |
| Human validation | Unsupported | No completed annotations |