BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence
Organizations: University of Illinois Urbana-Champaign
Abstract
Adaptive RAG uses signals such as confidence, relevance, support, and retrieval quality to decide when to search or correct evidence. In multi-step retrieval, however, these local signals must be combined into a persistent view of what the current evidence supports, what remains missing, and which action should follow. Existing methods often use such signals as separate triggers, making it difficult to preserve a coherent evidence state across a trajectory; we call this problem evidence-state fragmentation. We introduce BELIEFRAG, a closed-loop controller that updates an explicit state over sufficiency, reliability, conflict, uncertainty, evidence gaps, and acquisition cost, then chooses among retrieval, query rewriting, verification, answering, stopping, and abstention. Across six QA benchmarks with gpt-oss-120b, BELIEFRAG reaches mean token F1 0.572 with 3.89k tokens per question, outperforming fixed iterative retrieval (0.555 F1) while using 39% fewer tokens. The same quality-cost pattern transfers to Qwen3-32B, where BELIEFRAG reaches 0.552 F1 versus 0.523 for iterative retrieval while using 35% fewer tokens. Analysis shows that the main gains come from corrective re-retrieval rather than pruning alone, while several belief dimensions are redundant and calibrated answerability plays the strongest operational role. Calibration improves threshold stability across related evidence sources, although source shift can still invalidate the same decision signal.
Figures & tables
| Meaning | Computation | |
|---|---|---|
| Relevance | Softmax-weighted mean of calibrated top- retrieval scores. | |
| Support | Verifier score for how strongly supports a complete answer. | |
| Conflict | Largest material contradiction within or between evidence and the current draft. | |
| Uncertainty | Verifier estimate of how uncertain the answer remains given only . | |
| Gap | Estimated fraction of information required by the question that is still unsupported. | |
| Novelty | ; high when the newest retrieval adds information not already retained. |
| Decision | Condition | Default |
|---|---|---|
| Correct | Evidence is materially conflicting or clearly unreliable | or |
| Answer | Current state is operationally answerable and conflict is acceptable | |
| Retrieve | Evidence is insufficient, budget remains, and another retrieval has enough chance to make it answerable | |
| Rewrite | Another retrieval is useful, but the previous retrieval added little new information | |
| Stop / Abstain | Evidence is still insufficient and further acquisition has low expected value | no useful acquisition |
| Method | Multi-hop QA | Open-domain QA | Mean | Tokens | ||||
|---|---|---|---|---|---|---|---|---|
| HotpotQA | 2Wiki | MuSiQue | NQ | TriviaQA | PopQA | |||
| No-RAG ( Brown et al., 2020 ) | .378 | .411 | .167 | .341 | .766 | .381 | .407 | 0.10k |
| Static RAG ( Lewis et al., 2020 ) | .559 | .681 | .423 | .413 | .681 | .324 | .514 | 2.82k |
| Adaptive- ( Taguchi et al., 2025 ) | .535 | .558 | .360 | .347 | .715 | .348 | .477 | 3.04k |
| Adaptive-RAG ( Jeong et al., 2024 ) | .575 | .711 | .445 | .425 | .751 | .324 | .538 | 3.59k |
| Iterative RAG ( Jiang et al., 2023 ) | .588 | .710 | .420 | .420 | .750 | .440 | .555 | 6.35k |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Non-zero weights in | |||
|---|---|---|---|---|
| Sufficiency | 0.35 | 0.6 | ||
| Reliability | 0.50 | 0.5 | ||
| Conflict | 0.05 | 0.6 | ||
| Uncertainty | 0.70 | 0.6 | ||
| Gap | 0.90 | 0.7 | ||
| Cost | 0.00 | — | 1.0 | observed directly: |
| Step | Action | Decision and query/evidence change | |
|---|---|---|---|
| 0 | Retrieve | Initial belief is . With no evidence, the structural diagnostics set uncertainty and gap to . The fallback exceeds the retrieval threshold. Under first_query_is_question , the original question is issued directly as the first query. Five passages are returned; only Private_Music is useful. | 5 |
| 1 | Verify | The verifier reports approximately , , , and , with . Four passages are flagged unhelpful: Myron__duo_ , Zak_Starkey , Weathermaker_Music , and Oasis_discography . Verification prunes them, leaving only Private_Music . | 1 |
| 2 | Rewrite | The query is reformulated as Private Music signed drummer formerly a member of an English band . Rewriting changes the query but performs no retrieval. | 1 |
| 3 | Retrieve | The evidence-aware query writer reads the retained Private_Music passage, which names Ringo Starr , and makes that bridge entity explicit: Ringo Starr drummer member of which English group?> . Three additional passages are returned, but all are distractors and none mentions Ringo Starr or the Beatles. | 4 |
| 4 | Answer | With approximately , , , , conflict , and , the answer threshold is crossed. The final answer is The Beatles . | 4 |
| Budget | Value |
|---|---|
| Maximum retrieval rounds | 3 |
| Passages per retrieval (top- ) | 5 |
| Maximum controller actions | 6 |
| Token budget | 12,000 |
| Answer evidence window | 10 passages |
| Method | HotpotQA | 2Wiki | MuSiQue | NQ | TriviaQA | PopQA | Mean F1 |
|---|---|---|---|---|---|---|---|
| No-RAG | .265/.180/.320 | .382/.310/.420 | .154/.080/.200 | .210/.120/.280 | .512/.420/.550 | .185/.130/.220 | .285 |
| Static | .497/.380/.560 | .643/.570/.630 | .309/.170/.360 | .365/.230/.420 | .642/.560/.680 | .284/.240/.350 | .457 |
| Iterative | .559/.460/.610 | .659/.570/.630 | .361/.240/.410 | .412/.310/.470 | .695 /.600/.720 | .452/.370/.510 | .523 |
| Self-RAG* | .367/.300/.400 | .625/.517/.552 | .285/.210/.320 | .215/.110/.260 | .088/.080/.120 | .245/.210/.300 | .304 |
| CRAG | .512/.410/.570 | .680/.590/.660 | .352/.230/.400 | .388/.250/.440 | .665/.580/.700 | .410/.340/.460 | .501 |
| RL-Search | .278/.180/.310 | .285/.100/.310 | .198/.100/.230 | .092/.010/.120 | .248/.180/.290 | .210/.110/.250 | .219 |
| Method | HotpotQA | 2Wiki | MuSiQue | NQ | TriviaQA | PopQA | Mean EM/Judge |
|---|---|---|---|---|---|---|---|
| Static | .450/.640 | .600/.720 | .290/.440 | .280/.510 | .600/.810 | .280/.330 | .417/.575 |
| Iterative | .452/.660 | .670/.750 | .355/.520 | .350/.620 | .625/.860 | .410/.530 | .477/.657 |
| CRAG | .460/.650 | .630/.770 | .310/.460 | .280/.570 | .630/.830 | .380/.470 | .448/.625 |
| Self-RAG* | .380/.540 | .600/.680 | .270/.400 | .130/.310 | .100/.120 | .250/.290 | .288/.390 |
| RL-Search | .200/.620 | .110/.750 | .120/.440 | .010/.440 | .200/.530 | .130/.420 | .128/.533 |
| BeliefRAG | .480/.630 | .660/.810 | .370/.520 | .290/.580 | .650/.850 | .340/.430 | .465/.637 |