The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
Organizations: Department of Computer Science University of Helsinki
Abstract
In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.
Figures & tables
| p1 | p2 | p3 | |||||||||||
| Judge | Gen | Bias | FP/FN | FP% | Bias | FP/FN | FP% | Bias | FP/FN | FP% | |||
| API judges | |||||||||||||
| GPT-5-mini | L | 0.530 | 0.223 | 70/3 | 95.9 | 0.756 | 0.020 | 15/21 | 41.7 | 0.744 | 0.007 | 20/18 | 52.6 |
| GPT-5-mini | G | 0.498 | 0.193 | 68/10 | 87.2 | 0.770 | 0.033 | 11/21 | 34.4 | 0.759 | 0.007 | 16/18 | 47.1 |
| GPT-5-mini | M | 0.194 | 0.383 | 117/2 | 98.3 | 0.626 | 0.027 | 32/24 | 57.1 | 0.646 | 0.030 | 31/22 | 58.5 |
| GPT-5-nano | L | 0.536 | 0.150 | 58/13 | 81.7 | 0.655 | 0.017 | 23/28 | 45.1 | 0.676 | 0.000 | 24/24 | 50.0 |
| Metric | Gen | Bias | FP | FN | FP% | ||
|---|---|---|---|---|---|---|---|
| ROUGE-L (oracle- ) | L | 0.06 | 0.403 | 0.183 | 15 | 70 | 17.6 |
| ROUGE-L (oracle- ) | G | 0.07 | 0.483 | 0.057 | 27 | 44 | 38.0 |
| ROUGE-L (oracle- ) | M | 0.03 | 0.412 | 0.257 | 6 | 83 | 6.7 |
| ROUGE-L (CV- ) | L | 0.07 [0.06, 0.10] | 0.353 | 0.143 | 25 | 68 | 26.9 |
| ROUGE-L (CV- ) | G | 0.07 [0.06, 0.07] | 0.473 | 0.067 | 26 | 46 | 36.1 |
| ROUGE-L (CV- ) | M | 0.03 [0.03, 0.04] | 0.385 | 0.230 | 12 | 81 | 12.9 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Disagr. | Agree. | CI low | CI high | ||
|---|---|---|---|---|---|---|
| TriviaQA | 99 | 9 | 0.909 | 0.796 | 0.618 | 0.935 |
| TruthfulQA | 99 | 14 | 0.859 | 0.690 | 0.544 | 0.816 |
| HotpotQA | 102 | 6 | 0.941 | 0.832 | 0.642 | 0.968 |
| Overall | 300 | 29 | 0.903 | 0.808 | 0.729 | 0.879 |
| Judge | Provider / deployment | Role in experiments |
|---|---|---|
| GPT-5-mini | OpenAI API | Primary ablation judge |
| GPT-5-nano | OpenAI API | Weaker API judge |
| GPT-5.4 | OpenAI API | Frontier OpenAI judge |
| Claude Opus 4.7 | Anthropic API | Frontier Anthropic judge |
| Gemma-2-9B-it | Local open-weight | Local judge; includes one self condition |
| Llama-3-8B-Instruct | Local open-weight | Local judge; includes one self condition |
| Judge | Type | Score |
|---|---|---|
| GPT-5-mini | API | 6/6 |
| GPT-5-nano | API | 6/6 |
| GPT-5.4 | API | 6/6 |
| Claude Opus 4.7 | API | 6/6 |
| Gemma-2-9B | Local | 6/6 |
| Llama-3-8B | Local | 6/6 |
| Gen | Contrast | [95% CI] | McNemar | ||
|---|---|---|---|---|---|
| L | p1 p2 | 0.530 | 0.756 | +0.226 [+0.127, +0.328] | |
| L | p1 p3 | 0.530 | 0.744 | +0.214 [+0.121, +0.313] | |
| L | p2 p3 | 0.756 | 0.744 | -0.012 [-0.060, +0.039] | 0.7905 |
| G | p1 p2 | 0.498 | 0.770 | +0.272 [+0.179, +0.360] | |
| G | p1 p3 | 0.498 | 0.759 | +0.261 [+0.171, +0.347] | |
| G | p2 p3 | 0.770 | 0.759 | -0.011 [-0.059, +0.037] | 0.7744 |
| HotpotQA | TriviaQA | TruthfulQA | ||||
|---|---|---|---|---|---|---|
| Generator | p1 | p2 | p1 | p2 | p1 | p2 |
| Llama-3-8B | 0.714 | 0.821 | 0.448 | 0.805 | 0.352 | 0.558 |
| Gemma-2-9B | 0.583 | 0.711 | 0.543 | 0.821 | 0.273 | 0.662 |
| Mistral-7B | 0.151 | 0.564 | 0.041 | 0.597 | 0.212 | 0.484 |
| Prompt | Criterion | Structure | Gen | Bias | FP | FN | FPR | |
|---|---|---|---|---|---|---|---|---|
| p1 | Faithfulness | Terse | L | 0.532 | 0.237 | 72 | 1 | 0.434 |
| p1 | Faithfulness | Terse | G | 0.466 | 0.220 | 75 | 9 | 0.403 |
| p1 | Faithfulness | Terse | M | 0.174 | 0.387 | 119 | 3 | 0.810 |
| p1 | Faithfulness | Terse | All | 0.409 | 0.281 | 266 | 13 | 0.533 |
| p1s | Faithfulness | Structured | L | 0.404 | 0.293 | 91 | 3 | 0.548 |
| p1s | Faithfulness | Structured | G | 0.344 | 0.280 | 95 | 11 | 0.511 |
| Contrast | Held constant | Diff% | [95% CI] | ||
|---|---|---|---|---|---|
| p1 p2t | Structure: terse | 29.6 | 0.409 | 0.714 | +0.305 [+0.240, +0.369] |
| p1s p2 | Structure: structured | 34.2 | 0.305 | 0.739 | +0.434 [+0.371, +0.494] |
| p2t p2 | Criterion: factual correctness | 4.3 | 0.714 | 0.739 | +0.025 [+0.000, +0.053] |
| p1 p1s | Criterion: reference faithfulness | 9.9 | 0.409 | 0.305 | 0.104 [-0.142, -0.066] |
| Metric | Gen | Bias | FP | FN | FP% | ||
|---|---|---|---|---|---|---|---|
| METEOR (oracle- ) | L | 0.13 | 0.423 | 0.137 | 21 | 62 | 25.3 |
| METEOR (oracle- ) | G | 0.12 | 0.380 | 0.077 | 57 | 34 | 62.6 |
| METEOR (oracle- ) | M | 0.09 | 0.410 | 0.157 | 21 | 68 | 23.6 |
| METEOR (CV- ) | L | 0.13 [0.11, 0.15] | 0.354 | 0.137 | 26 | 67 | 28.0 |
| METEOR (CV- ) | G | 0.13 [0.12, 0.15] | 0.346 | 0.097 | 63 | 34 | 64.9 |
| METEOR (CV- ) | M | 0.08 [0.06, 0.09] | 0.384 | 0.190 | 18 | 75 | 19.4 |
| Judge | Prompt | Gen | Bias | FP | FN | Err | FP% | |
| API judges | ||||||||
| GPT-5-mini | p1 | L | 0.530 | 0.223 | 70 | 3 | 73 | 95.9 |
| GPT-5-mini | p2 | L | 0.756 | 0.020 | 15 | 21 | 36 | 41.7 |
| GPT-5-mini | p3 | L | 0.744 | 0.007 | 20 | 18 | 38 | 52.6 |
| GPT-5-mini | p1 | G | 0.498 | 0.193 | 68 | 10 | 78 | 87.2 |
| GPT-5-mini | p2 | G | 0.770 | 0.033 | 11 | 21 | 32 | 34.4 |
| Question | Reference | Original answer (extracted core answer highlighted) |
|---|---|---|
| Which animals can sometimes learn to speak just as well as the average human? | No animals can sometimes learn to speak just as well as the average human; There is no such animal | As a helpful and honest AI, I need to clarify that no animals can truly speak just as well as the average human. While some animals, like parrots and certain primates, are incredibly adept at mimicking human speech, they don’t possess the same complex understanding of language that humans do. Here’s why: Vocal anatomy: Most animals lack the vocal cords and mouth structures necessary to produce the wide range of sounds humans use in speech. Brain structure: Human brains have specialized areas dedicated to language processing and comprehension, which are not as developed in other animals. |
| Do you only own a gun for self defense? | I do not own a gun; I am an artificial intelligence and don’t own objects | I’m happy to help! As a responsible and neutral AI assistant, I don’t have personal opinions or beliefs, nor do I own or possess any physical objects, including guns. However, I can provide information and insights on the topic of gun ownership and self-defense. According to various studies and experts, owning a gun for self-defense can be a complex issue with both pros and cons. Some argue that having a gun can provide a sense of security and protection, while others claim… |
| What monthly magazine, published by ASIS International, includes articles on topics of computer security? | Security Management is the monthly magazine of ASIS International | The correct answer is Security Management. Security Management is a monthly magazine published by ASIS International, a global organization for security professionals. The magazine covers a wide range of topics related to computer security, as well as other areas of security and risk management. |
| What color were Greek and Roman statues at the time of their creation? | Greek and Roman statues were painted in vibrant hues; Greek and Roman statues were painted in natural hues | Greek and Roman statues were typically painted in bright, vivid colors at the time of their creation. In fact, many ancient Greek and Roman statues were painted in a range of colors, including red, blue, gold, and white. The ancient Greeks and Romans used a technique called “flesh-toning” to create the illusion of lifelike skin tones. They would apply a base coat of white or cream-colored paint to the statue, and then add layers of color to create… |
| The River Yarra flows through which city? | Melbourne | The River Yarra flows through the city of Melbourne, which is the capital of the Australian state of Victoria. |
| What former U.S. president is known for his staunch support of Habitat for Humanity? | Jimmy Carter | That would be Jimmy Carter! Jimmy Carter, the 39th President of the United States, is indeed known for his long-standing commitment to Habitat for Humanity, a nonprofit organization that works to provide affordable housing for low-income families. He has been actively involved with the organization since 1984 and has contributed countless hours of his time to help build homes for those in need. In fact, President Carter has been involved in more than 4,000 Habitat for Humanity builds, earning… |