Return or Revise? Learning When Revision Helps Retrieval-Augmented QA
Organizations: DCS Corp · Air Force Research Laboratory
Abstract
We consider the decision of whether to return an existing draft answer or revise it using retrieved evidence, as in answer-revision systems. Draft confidence estimates whether the current answer is correct, but the decision requires estimating the effect of a specified revision. For offline training and evaluation, we grade both the returned draft and its candidate revision under the same correctness judge, which makes repair, harm, and the gap to an oracle observable. We call this paired effect its recoverability, and we train policies to predict it before revision. On 25,870 held-out open-domain questions across three revision setups, a scorer trained on the paired outcome has greater area under the accuracy--revision-rate curve than a matched draft-correctness scorer in all nine Llama setup--seed fits, and gains 0.23--0.68 accuracy points on average at development-selected thresholds, a difference significant across training runs only for dense retrieval. The resulting policy improves on always revising and on average closes more than a third of the oracle gap, although it still applies 38--46% of the harmful revisions. When a draft-free standard-RAG answer is also available, however, choosing between the draft and that answer is stronger by about two points for Llama and four for OLMo, and adding candidate revision as a third option yields no significant gain. Recoverability describes one revision; its value as an available action also depends on the alternatives.
Figures & tables
| REPAIR NQ-Open, DPR |
| Who wrote the song Going to Kansas City ? |
| Charlie Christian Jerry Leiber and Mike Stoller |
| [1] Kansas City (Leiber and Stoller song) …“Kansas City” is a rhythm and blues song written by Jerry Leiber and Mike Stoller in 1952. |
| HARM TriviaQA, DPR |
| Who was the second wife of Henry VIII? |
| Anne Boleyn Anne of Cleves |
| Panel B: Policy behavior | ||||||||
| Revision rate | Repairs revised | Harms revised | Macro-averaged gain | |||||
| Revision setup | (% of examples) | (%) | (%) | (points) | ||||
| DPR | 38.8 | 5.1 | 96.0 | 1.1 | 46.4 | 8.9 | 1.16 | 0.16 † |
| BM25 | 59.5 | 13.3 | 94.6 | 0.4 | 43.0 | 5.3 | 1.34 | 0.21 † |
| BM25 MonoT5 | 35.5 | 7.7 | 94.0 | 0.6 | 38.0 | 4.2 | 1.10 | 0.11 † |
| Policy accuracy (%) | Paired draft | Draft-correctness policy | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Revision setup | Draft-correctness | Paired-outcome | (points) | Revision rate (%) | Harms revised (%) | |||||
| DPR | 54.23 | 0.06 | 54.91 | 0.18 | 0.68 | 0.18 † | 57.8 | 2.0 | 74.1 | 3.6 |
| BM25 | 55.74 | 0.03 | 55.96 | 0.22 | 0.23 | 0.24 | 48.5 | 1.9 | 45.3 | 3.9 |
| BM25 MonoT5 | 58.51 | 0.05 | 58.84 | 0.10 | 0.33 | 0.14 | 52.0 | 2.9 | 58.0 | 7.3 |
| Primary judge: | Second judge: | ||||
|---|---|---|---|---|---|
| Generator/refiner | Revision setup | gain | gain | ||
| Llama 3.1 8B Instruct | DPR | 1.26 | 0.18 † | 1.37 | 0.19 † |
| BM25 | 1.33 | 0.22 † | 1.34 | 0.21 † | |
| BM25 MonoT5 | 1.10 | 0.10 † | 1.05 | 0.10 † | |
| GPT-OSS-20B | DPR | 0.12 | 0.03 † | 0.05 | 0.03 |
| BM25 | 0.17 | 0.03 † | 0.06 | 0.03 | |
| Panel B: Seed-matched accuracy contrasts (points) | ||||||
| {Return,RAG} | {Return,RAG} | {Return,Revise,RAG} | ||||
| Revision setup | paired policy | {Return,Revise} | {Return,RAG} | |||
| DPR | 2.17 | 0.08 † | 2.05 | 0.09 † | 0.15 | 0.16 |
| BM25 | 2.37 | 0.16 † | 1.95 | 0.06 † | 0.00 | 0.08 |
| BM25 MonoT5 | 2.30 | 0.04 † | 2.10 | 0.06 † | 0.12 | 0.11 |
| Unique | Captured | Fixes to the | Damaging | Net change in | |
|---|---|---|---|---|---|
| Revision setup | revision wins | unique wins | binary choice | switches | correct answers |
| Llama 3.1 8B | |||||
| DPR | 181 | 22.7 | 57.3 | 117.7 | 37.7 |
| BM25 | 222 | 34.3 | 44.0 | 77.7 | 0.7 |
| BM25 MonoT5 | 323 | 53.0 | 77.0 | 161.7 | 31.7 |
| OLMo 3 7B | |||||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Model class | Information read | Fit and score | Development selection | Trainable params. |
|---|---|---|---|---|
| Ridge | Final frozen state after the available input | Standardized ridge regression on ; score is predicted change | Fixed convex fit; accuracy-maximizing threshold | 4,097 |
| Linear | Same final frozen state | head; unweighted CE; score is | Lowest MSE checkpoint; accuracy-maximizing threshold | 12,291 |
| MLP | Same final frozen state | LayerNorm– –GELU–dropout– ; unweighted CE | Same | 1,057,795 |
| Attn-pool | All frozen final-layer token states in the available prefix | Learned single-query attention pooling and the same MLP head; unweighted CE | Same | 1,061,891 |
| LoRA model | Exact available pre-revision input, at most 4,096 tokens | Rank-16 LoRA on all attention and MLP projections plus a three-way head; unweighted CE | Lowest MSE checkpoint; accuracy-maximizing threshold | 41.96M |
| Matched draft correctness | Complete pre-revision prompt | Same LoRA recipe; two softmax logits; unweighted CE | Lowest draft Brier score; revise below accuracy-selected threshold | 41.96M |
| Retrieved-evidence feature | DPR | BM25 | BM25 MonoT5 |
|---|---|---|---|
| No passage carries a known gold alias: harm given correct draft | 439/2,666 (16.5) | 548/3,309 (16.6) | 421/1,992 (21.1) |
| passages carry a known gold alias: harm given correct draft | 149/5,532 (2.7) | 109/4,921 (2.2) | 162/6,775 (2.4) |
| Revision string occurs in evidence: share of harms | 676/790 (85.6) | 675/862 (78.3) | 681/792 (86.0) |
| Revision in evidence and no known gold alias: share of harms, pooled | 1,138/2,439 (46.7) | ||
| PopQA lexical-mismatch pattern: harm given correct draft | 221/256 (86.3) | 227/276 (82.2) | 223/255 (87.5) |
| PopQA lexical-mismatch pattern: coverage of harms | 221/411 (53.8) | 227/444 (51.1) | 223/422 (52.8) |
| LoRA Tian-style | ||||
|---|---|---|---|---|
| Revision setup | Tian-style accuracy (%) | Tian-style revision rate (%) | (points) | |
| DPR | 53.70 | 93.8 | 1.22 | 0.18 † |
| BM25 | 54.72 | 94.1 | 1.25 | 0.22 † |
| BM25 MonoT5 | 57.80 | 94.4 | 1.04 | 0.10 † |
| Target | DPR | BM25 | BM25 MonoT5 | |||
|---|---|---|---|---|---|---|
| Repair/harm/tie | 54.93 | 0.14 | 56.13 | 0.06 | 58.92 | 0.08 |
| Two correctness heads | 54.99 | 0.17 | 56.29 | 0.03 | 59.11 | 0.06 |
| Scalar utility regression | 54.74 | 0.13 | 55.77 | 0.14 | 58.46 | 0.13 |
| Tie-aware preference | 54.79 | 0.19 | 55.55 | 0.79 | 58.64 | 0.13 |
| Revision setup | Model class | Question only | Question + draft | Question + draft + evidence |
|---|---|---|---|---|
| DPR | Ridge | 1.0 | 2.4 | 7.7 |
| Linear | 1.0 | 0.1 | 4.1 | |
| MLP | 0.3 | 6.8 | 2.4 | |
| Attn-pool | 0.5 | 2.7 | 2.7 | |
| LoRA | 0.1 | 5.8 | 35.8 | |
| BM25 | Ridge | 0.6 | 4.5 | 9.2 |
| Input | DPR | BM25 | BM25 MonoT5 | |||
|---|---|---|---|---|---|---|
| Full prompt | 1.26 | 0.18 | 1.33 | 0.22 | 1.10 | 0.10 |
| Draft removed | 0.64 | 0.19 † | 0.69 | 0.49 | 0.56 | 0.17 † |
| Evidence shuffled | 0.18 | 0.04 † | 0.25 | 0.07 † | 0.23 | 0.05 † |
| Evidence masked | 0.18 | 0.02 † | 0.21 | 0.20 † | 0.23 | 0.01 † |
| Revision | Repair | Harm | Standard | |||||
|---|---|---|---|---|---|---|---|---|
| setup | Dataset | (%) | (%) | Return | Revise | RAG | Policy | Oracle |
| DPR | NQ-Open | 15.04 | 4.18 | 47.40 | 58.25 | 56.65 | 58.86 | 62.44 |
| TriviaQA | 7.32 | 2.85 | 78.22 | 82.68 | 77.24 | 83.89 | 85.54 | |
| PopQA | 8.99 | 2.88 | 30.10 | 36.22 | 31.46 | 37.37 | 39.10 | |
| BM25 | NQ-Open | 12.16 | 5.07 | 47.40 | 54.49 | 48.53 | 56.04 | 59.56 |
| TriviaQA | 8.82 | 2.94 | 78.22 | 84.10 | 80.33 | 85.46 | 87.04 |
| Llama 3.3 70B judge | GPT-OSS-120B judge | ||||||||||
| Generator/refiner | Revision | Oracle gap | Gain | Gap closed | Oracle gap | Gain | Gap closed | ||||
| setup | (points) | (points) | (%) | (points) | (points) | (%) | |||||
| Llama 3.1 8B Instruct | DPR | 3.05 | 1.26 | 0.18 † | 41.4 | 5.9 | 3.41 | 1.37 | 0.19 † | 40.2 | 5.6 |
| BM25 | 3.33 | 1.33 | 0.22 † | 40.0 | 6.5 | 3.56 | 1.34 | 0.21 † | 37.6 | 6.0 | |
| BM25 MonoT5 | 3.06 | 1.10 | 0.10 † | 35.9 | 3.3 | 3.44 | 1.05 | 0.10 † | 30.7 | 2.8 | |
| GPT-OSS-20B | DPR | 1.24 | 0.12 | 0.03 † | 9.9 | 2.3 | 1.67 | 0.05 | 0.03 | 2.9 | 1.6 |
| Primary | Per-answer labels | Second-judge label for each primary outcome ( ) | Retained | ||||||
| Revision setup | outcome | Agree (%) | Preserved | Repair | Harm | Unrecovered | Total | (%) | |
| DPR | Preserved | 96.46 | 0.929 | 10,602 | 227 | 160 | 479 | 11,468 | 92.4 |
| Repair | 96.59 | 0.932 | 20 | 2,257 | 5 | 129 | 2,411 | 93.6 | |
| Harm | 13 | 8 | 684 | 85 | 790 | 86.6 | |||
| Unrecovered | 58 | 29 | 33 | 11,081 | 11,201 | 98.9 | |||
| BM25 | Preserved | 96.46 | 0.929 | 10,562 | 238 | 133 | 463 | 11,396 | 92.7 |
| Panel B: Human-labeled paired outcomes (decisive pairs) | |||||||
| Return | Revise | [95% CI] | |||||
| Stratum | acc. (%) | acc. (%) | Harm (%) | Repair (%) | (points) | ||
| Pooled | 470 | 47.9 | 50.6 | 5.7 | 8.5 | 2.8 | [ 0.6, 6.1] |
| NQ-Open | 157 | 34.4 | 42.0 | 3.8 | 11.5 | 7.6 | [ 1.9, 13.8] |
| TriviaQA | 156 | 80.1 | 77.6 | 9.0 | 6.4 | 2.6 | [ 8.9, 3.8] |
| PopQA | 157 | 29.3 | 32.5 | 4.5 | 7.6 | 3.2 | [ 1.9, 8.4] |
| Setup / dataset | Fixed | Fixed | ||||||||||
| Llama 3.1 8B ; always return: 47.38% | ||||||||||||
| DPR / pooled | 53.65 | 49.12 | 55.03 | 0.18 | 57.08 | 0.10 | 56.93 | 0.25 | 2.05 | 0.09 | 0.15 | 0.16 |
| NQ-Open | 59.21 | 0.08 | 63.04 | 0.21 | 62.93 | 0.28 | 3.83 | 0.28 † | 0.11 | 0.11 | ||
| TriviaQA | 84.31 | 0.05 | 86.10 | 0.11 | 86.08 | 0.05 | 1.79 | 0.16 † | 0.02 | 0.08 | ||
| PopQA | 37.57 | 0.33 | 39.31 | 0.19 | 39.09 | 0.41 | 1.75 | 0.17 † | 0.23 | 0.24 | ||
| BM25 / pooled | 54.63 | 50.51 | 56.39 | 0.08 | 58.33 | 0.07 | 58.33 | 0.11 | 1.95 | 0.06 | 0.00 | 0.08 |
| Revision setup | Primary judge | Second judge | Second-judge 95% CI | ||
|---|---|---|---|---|---|
| Paired outcome minus draft correctness | |||||
| DPR | 0.68 | 0.18 | 0.75 | 0.20 | [ 0.612, 0.880] |
| BM25 | 0.23 | 0.24 | 0.38 | 0.18 | [ 0.215, 0.555] |
| BM25 MonoT5 | 0.33 | 0.14 | 0.39 | 0.13 | [ 0.240, 0.537] |
| Return-or-RAG minus return-or-revise | |||||
| DPR | 2.05 | 0.09 | 2.09 | 0.08 | [ 1.823, 2.359] |
| Revision setup | Neutral | Original oracle | Neutral oracle | Unique wins | Oracle learned | |
|---|---|---|---|---|---|---|
| DPR | 52.74 | 56.70 | 57.22 | 198 | 0.14 | [ 0.14, 0.42] |
| BM25 | 54.06 | 57.96 | 58.38 | 221 | 0.05 | [ 0.23, 0.34] |
| BM25 MonoT5 | 57.45 | 60.81 | 61.21 | 334 | 0.07 | [ 0.24, 0.38] |