RAIM: Robust Aggregation of Inexpensive Models for Hallucination Detection
Organizations: Computer, Electrical and Mathematical Sciences and Engineering (CEMSE) Division, King Abdullah University of Science and Technology (KAUST) Thuwal 23955, Saudi Arabia
Abstract
Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary frontier models, costly and ill-suited to high-throughput monitoring. We investigate whether a panel of cheap open-weight judges (4--9B) can be aggregated to stand in for a frontier one, what the substitution sacrifices, and when it is worth making. We propose RAIM, an aggregation scheme robust to the members' correlated errors, coupling a cross-fitted stacked logistic regression with an admissibility test that, read from the members' own outputs, identifies when aggregating them improves on their best member and stays within reach of the frontier judge. We instantiate RAIM with ten judges from disjoint families across eight faithfulness benchmarks. Against Claude Sonnet, the panel retains a median 93% of its Cohen's and gives up only 2.9 points of balanced accuracy on average; read as paired differences, it clearly improves on one benchmark and clearly worsens on three (only two by a non-negligible margin), leaving four unresolved. At a sixty-fourth of the frontier's inference price, the operative expense is a one-time in-domain calibration on 50--100 labelled records. The panel is also competitive with purpose-trained detectors on their home benchmarks (within 1.3 accuracy points of GPT-4o and 1.9 of the LLM-AggreFact leader), and beats the strongest one we reran by 6 points on our grounded sets. Whether aggregation pays depends on the members themselves: where several capable members err on different items, the panel improves on its best judge and approaches the frontier; where one dominates, the stacker recovers the leader, and only there does the frontier remain materially ahead. Both conditions are read off the calibration set at no further cost, so a cheap panel can stand in for a frontier one wherever this audit admits it.
Figures & tables
| label-free | single judge, label-selected | stacked panel | frontier ( Sonnet ) | paired (panel comparator) | |||||||||||
| Dataset | avg. mem. | unw. | D–S | best (avg) | best (CV) | oracle | bacc | bacc | avg ( ) | best ( ) | Sonnet ( ) | Sonnet ( bacc ) | |||
| Core grounded | |||||||||||||||
| MedHallu | 2000 | 0.590 | 0.708 | 0.716 | 0.656 | 0.693 † | 0.693 | 0.725 | 86.3 | 0.761 | 88.1 | ‡ | |||
| WiCE | 222 | 0.385 | 0.449 | 0.486 | 0.547 | 0.484 † | 0.547 | 0.540 | 77.0 | 0.548 | 77.4 | ||||
| RAGTruth | 1362 | 0.181 | 0.179 | 0.325 | 0.399 | 0.399 | 0.399 | 0.409 | 70.4 | 0.601 | 80.0 | ||||
| XSum | 546 | 0.312 | 0.395 | 0.420 | 0.377 | 0.400 † | 0.437 | 0.437 | 71.9 | 0.418 | 70.9 | ||||
| quality ( bacc ) by scope | cost ( usd ) | scope | ||||||
| Route | size | LLM-AggreFact -4 | grounded 6 | suite 8 | per 1k | one-time | open | ungr. |
| purpose-trained and prompted comparators | ||||||||
| Bespoke-MiniCheck-7B (leaderboard) | 7B | — | n/a | fine-tune | ||||
| MiniCheck-Flan-T5-L [ 73 ] | 0.8B | n/a | fine-tune | |||||
| Granite Guardian 3.3 (leaderboard) | 8B | — | — | fine-tune | ||||
| Prometheus-2 [ 39 ] | 7B | fine-tune | ||||||
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| App. | What it carries | Supports |
| The account, its extensions, and its bounds | ||
| B | The competence/correlation account in full: the item-level recovery test that carries it, the cross-dataset predictor that does not, the regime ordering read dataset by dataset, the member-level views, the panel-size sweep, and the effective-independence comparison against a published frontier panel | § 3.3 |
| C | The identity that makes the correlation coordinate free, the obstruction that keeps the competence one from being so, the threshold family and its check on (sub-)panels, the labelling budget for both the test and the aggregator, and which half of the rule these eight datasets actually test | § 3.3 , § 4 |
| D | Matched base-versus-instruct probe: how instruction tuning moves the panel’s error-correlation | § 3.3 |
| E | Leave-one-dataset-out transfer, and why the in-domain advantage does not survive it | § 3.3 |
| External comparisons and their calibration | ||
| members correct | ||
| 0 | 0.000 | 529 |
| 1 | 0.058 | 327 |
| 2 | 0.162 | 235 |
| 3 | 0.101 | 189 |
| 4 | 0.157 | 210 |
| 5 | 0.225 | 160 |
| Spearman (two-sided ) against the (stacked best-single) gap | ||||
| subset | members | members | ||
| all eight | 8 | (0.39) | (0.23) | (0.24) |
| without CNN | 7 | (0.15) | (0.06) | (0.07) |
| without CNN and ExpertQA | 6 | (0.47) | (0.23) | (0.27) |
| full sample | (call matches) at a budget of | ||||||||
| Dataset | spread | regime | 10 | 25 | 50 | 100 | 200 | ||
| MedHallu | 2000 | admissible | |||||||
| WiCE | 222 | admissible | |||||||
| RAGTruth | 1362 | not admissible | |||||||
| XSum | 546 | not admissible | |||||||
| CNN | 114 | not admissible | — | ||||||
| Dataset | train | unw. | labelled records the stacker is fitted on | |||||||
| 10 | 25 | 50 | 100 | 200 | 400 | 800 | full | |||
| MedHallu | 800 | |||||||||
| WiCE | 176 | — | — | — | ||||||
| RAGTruth | 544 | — | ||||||||
| XSum | 377 | — | — | |||||||
| CNN | 87 | — | — | — | — | |||||
| System | size | CNN | XSum | WiCE | ExpertQA | mean | source |
| published, from the leaderboard and the cited papers | |||||||
| Bespoke-MiniCheck-7B | 7B | leaderboard | |||||
| MiniCheck-Flan-T5-L | 0.8B | [ 73 ] | |||||
| FactCG-DeBERTa-L | 0.4B | [ 45 ] , leaderboard | |||||
| HalluGuard-4B † | 4B | [ 7 ] | |||||
| Granite Guardian 3.3 | 8B | leaderboard | |||||
| System | F1 | |
| MedHallu — published | ||
| GPT-4o [ 59 ] | frontier, prompted | |
| Qwen2.5-14B [ 59 ] | single open judge | |
| GPT-4o-mini [ 59 ] | budget api | |
| Qwen2.5-7B [ 59 ] | our panel member | |
| Gemma-2-9B [ 59 ] | our panel member | |
| published anchor | ours ( bacc ) | |||||
| Dataset | strongest reported system | value | metric | what blocks a direct comparison | panel | Sonnet |
| TruthfulQA | GPT-judge (fine-tuned GPT-3-6.7B ) [ 50 ] | – | accuracy | judges model generations after fine-tuning on in-domain human ratings; we score the dataset’s own reference answers zero-shot | ||
| FActScore | retrieval-augmented estimator [ 55 ] | — | system-level error | retrieves Wikipedia evidence by construction, approximating the human score to within at corpus level; we supply no evidence and score per claim | ||
| unweighted panel bacc | members at | Sonnet | |||||
| Dataset | def_on | claim | def_on | claim | def_on | claim | |
| XSum | [+14.5, +21.7] | ||||||
| WiCE | [+10.0, +22.0] | ||||||
| CNN | [+0.8, +7.0] | ||||||
| ExpertQA | [+3.2, +7.7] | ||||||
| purpose-trained judge | panel | |||||||
| Dataset | Prometheus-2 | JudgeLM | Auto-J | MiniCheck- Flan-T5-L ⋄ | Bespoke- MiniCheck-7B ⋄ | stacked | best single (CV) | Sonnet |
| MedHallu | 0.448 | -0.204 | 0.003 | 0.262 | -0.010 | 0.725 | 0.693 | 0.761 |
| WiCE | 0.224 | -0.009 | 0.002 | 0.459 | 0.214 | 0.540 | 0.484 | 0.548 |
| RAGTruth | 0.004 | 0.028 | -0.027 | 0.262 | 0.338 | 0.409 | 0.399 | 0.601 |
| XSum | 0.169 | 0.051 | 0.053 | 0.484 | 0.267 | 0.437 | 0.400 | 0.418 |
| CNN | 0.226 | 0.052 | -0.000 | 0.349 | 0.103 | 0.349 | 0.298 | 0.469 |
| Cohen’s | (panel Sonnet ) | |||||
| Dataset | stacked | Qwen2.5-32B | Qwen2.5-72B | Sonnet | bacc | |
| MedHallu | 0.725 | 0.614 | 0.669 | 0.761 | * [-0.069, -0.003] | * [-3.5, -0.1] |
| WiCE | 0.540 | 0.522 | 0.629 | 0.548 | [-0.140, +0.123] | [-6.8, +6.1] |
| RAGTruth | 0.409 | 0.484 | 0.351 | 0.601 | * [-0.239, -0.144] | * [-12.0, -7.2] |
| XSum | 0.437 | 0.452 | 0.446 | 0.418 | [-0.069, +0.102] | [-3.3, +5.1] |
| CNN | 0.349 | 0.380 | 0.349 | 0.469 | [-0.275, +0.048] | [-13.8, +2.1] |
| Dataset | tier | frame | content | |
| MedHallu | 2000 | core grounded | QA | medical question answering with synthesised hallucinations [ 59 ] |
| WiCE | 222 | core grounded | claim | sentence-level entailment of Wikipedia claims against cited evidence [ 34 ] |
| RAGTruth | 1362 | core grounded | QA | retrieval-augmented responses annotated for unsupported content [ 58 ] |
| XSum | 546 | core grounded | claim | AggreFact factuality annotations on extreme-summarisation output [ 72 , 57 ] |
| CNN | 114 | grounded | claim | AggreFact factuality annotations on CNN/DailyMail summaries [ 72 ] |
| ExpertQA | 1462 | grounded | claim | expert-curated questions with attributed answers [ 52 ] |
| method | MedHallu | WiCE | RAGTruth | XSum | CNN | ExpertQA | FActScore | TruthfulQA |
| unweighted panel | 0.708 | 0.449 | 0.179 | 0.395 | 0.070 | 0.146 | 0.485 | 0.374 |
| accuracy-weighted | 0.710 | 0.468 | 0.303 | 0.395 | 0.053 | 0.190 | 0.485 | 0.411 |
| log-odds-weighted | 0.711 | 0.468 | 0.303 | 0.394 | 0.053 | 0.191 | 0.473 | 0.409 |
| average member | 0.590 | 0.385 | 0.181 | 0.312 | 0.088 | 0.154 | 0.259 | 0.274 |
| best-on-average single | 0.656 | 0.547 | 0.399 | 0.377 | 0.298 | 0.217 | 0.444 | 0.471 |
| best single (CV) | 0.693 | 0.484 | 0.399 | 0.400 | 0.298 | 0.240 | 0.444 | 0.471 |
| method | MedHallu | WiCE | RAGTruth | XSum | CNN | ExpertQA | FActScore | TruthfulQA |
| unweighted panel | 85.4 | 72.5 | 59.0 | 69.8 | 53.5 | 57.3 | 74.3 | 68.7 |
| accuracy-weighted | 85.5 | 73.4 | 65.2 | 69.7 | 52.6 | 59.5 | 74.3 | 70.5 |
| log-odds-weighted | 85.6 | 73.4 | 65.2 | 69.7 | 52.6 | 59.6 | 73.6 | 70.5 |
| average member | 77.2 | 63.9 | 58.6 | 65.0 | 54.2 | 55.2 | 62.6 | 63.5 |
| best-on-average single | 82.8 | 77.1 | 69.9 | 68.9 | 64.9 | 60.9 | 70.3 | 73.4 |
| best single (CV) | 70.0 | 73.9 | 69.9 | 69.9 | 64.9 | 62.0 | 70.3 | 73.4 |
| Cohen’s | rescue | unsupervised aggregator | |||||
| Dataset | unw. | D–S | stk. | ||||
| MedHallu | 0.708 | 0.716 | 0.725 | [+0.00, +0.03] | [-0.01, +0.04] | [-0.01, +0.03] | [-0.01, +0.05] |
| WiCE | 0.449 | 0.486 | 0.540 | [-0.01, +0.19] | [-0.14, +0.08] | [-0.04, +0.15] | [-0.10, +0.11] |
| RAGTruth | 0.179 | 0.325 | 0.409 | * [+0.18, +0.27] | * [-0.26, -0.18] | * [+0.04, +0.12] | * [-0.11, -0.04] |
| XSum | 0.395 | 0.420 | 0.437 | [-0.02, +0.10] | [-0.07, +0.06] | [-0.03, +0.06] | [-0.04, +0.08] |
| FActScore | 0.485 | 0.515 | 0.503 | [-0.05, +0.08] | [-0.06, +0.14] | [-0.05, +0.02] | [-0.01, +0.15] |
| Dataset | source | upstream pool | items | clusters | per cl. | balance |
| one matched pair per source record: a supported and a hallucinated candidate sharing its question and, on the grounded sets, its evidence | ||||||
| MedHallu | MedHallu/pqa_labeled [ 59 ] | questions | / | |||
| RAGTruth | RAGTruth-processed , QA [ 58 ] | resp. / q. | / | |||
| TruthfulQA | truthful_qa/generation [ 50 ] | questions | / | |||
| FActScore | InstructGPT.jsonl [ 55 ] | biographies | / | |||
| classes pooled across documents and the majority subsampled to the minority count (seed ), clusters recovered from the shared document | ||||||
| MedHallu — 0:pos and 0:neg , same cluster | |
| question | Do mitochondria play a role in remodelling lace plant leaves during programmed cell death? |
| evidence | Programmed cell death (PCD) is the regulated death of cells within an organism. The lace plant (Aponogeton madagascariensis) produces perforations in … |
| candidate, supported | Results depicted mitochondrial dynamics in vivo as PCD progresses within the lace plant, and highlight the correlation o … |
| candidate, hallucinated | Mitochondria regulate the formation of perforations in lace plant leaves through the modulation of calcium channels and … |
| RAGTruth — 0:pos and 0:neg , same cluster | |
| question | butcher shop phone number |
| WiCE — 1:pos and 0:neg , different clusters | |
| question | Each player received a key to the city from Mayor Bill de Blasio. |
| evidence | TITLE: Megan Rapinoe - #SheBelieves: historic ticker-tape parade in NYC for U.S. women’s national soccer team - Pictures - CBS News PUBLISHER: https:/ … |
| candidate, supported | Each player received a key to the city from Mayor Bill de Blasio. |
| candidate, hallucinated | The diocese is currently a titular see of the Patriarchate of Constantinople, and Gerasimos Papadopoulos was titular Bis … |
| XSum — 0:pos and 1:neg , different clusters | |
| question | The number of recorded homicides in Scotland has fallen to its lowest level for more than 40 years, according to the Sco … |
| # | tag | checkpoint | developer | ref. |
| 00 | Llama | meta-llama/Llama-3.1-8B-Instruct | Meta | [ 51 ] |
| 01 | Qwen | Qwen/Qwen2.5-7B-Instruct | Alibaba | [ 62 ] |
| 02 | Mistral | mistralai/Mistral-7B-Instruct-v0.3 | Mistral AI | [ 31 ] |
| 03 | Gemma | google/gemma-2-9b-it | [ 17 ] | |
| 04 | Phi | microsoft/Phi-4-mini-instruct | Microsoft | [ 54 ] |
| 05 | Yi | 01-ai/Yi-1.5-9B-Chat-16K | 01.AI | [ 1 ] |
| member | MedHallu | WiCE | RAGTruth | XSum | CNN | ExpertQA | FActScore | TruthfulQA |
| Llama | 1.000 | 0.806 | 1.000 | 0.982 | 0.956 | 0.966 | 0.997 | 0.998 |
| Qwen | 1.000 | 0.991 | 1.000 | 0.998 | 1.000 | 1.000 | 1.000 | 0.999 |
| Mistral | 1.000 | 0.986 | 0.996 | 0.996 | 1.000 | 0.989 | 1.000 | 1.000 |
| Gemma | 1.000 | 0.995 | 1.000 | 1.000 | 1.000 | 1.000 | 0.973 | 0.998 |
| Phi | 0.998 | 0.523 | 0.999 | 1.000 | 1.000 | 0.756 | 1.000 | 0.996 |
| Yi | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 0.997 | 1.000 | 1.000 |
| abstentions | stacked | ||||||
| Dataset | count | rate | categ. | (categ. ) | moved | ||
| MedHallu | 2000 | 567 | 12 | ||||
| RAGTruth | 1362 | 123 | 5 | ||||
| WiCE | 222 | 179 | 7 | ||||
| XSum | 546 | 53 | 8 | ||||
| CNN | 114 | 5 | 0 | ||||