CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation
Organizations: International Institute of Information Technology Bhubaneswar, India · Salesforce India Pvt Ltd, Bengaluru, India
Abstract
Faithfulness evaluation of abstractive summaries remains an open challenge, with existing metrics addressing only isolated hallucination types: factual entity errors, relational inconsistencies, or numerical fabrications, without capturing their co-occurrence or interaction. We introduce CHI (Composite Hallucination Index), the first unified hallucination metric that decomposes faithfulness errors into three orthogonal dimensions: entity hallucination (EHI), relation hallucination (RHI*), and quantity hallucination (QHI). Each dimension employs a shared softmax-normalized architecture over Venn diagram-derived factors representing extractiveness, positive hallucination, over-focus, negative hallucination, and lost focus. The novel QHI component introduces tolerance-aware numerical matching with exact, epsilon, derived, and temporal comparison modes. We fuse the three dimensions via harmonic mean to produce a single composite score that penalizes weakness in any dimension. We validate CHI on 800 source articles spanning four domains (news, medical, legal, financial) with summaries from five generation systems. Empirical results demonstrate that: (i) the three dimensions are statistically orthogonal (mean rho = 0.148), confirming they capture distinct error types; (ii) CHI achieves the highest system-level correlation with human judgments (rho = 0.66, p = 0.006) on SummEval, outperforming ROUGE (rho = 0.53), EHI (rho = 0.58), and all individual components; and (iii) ablation studies confirm that all three dimensions contribute unique variance, with the full composite outperforming any individual component while providing decomposable error diagnostics unavailable from single-score baselines. CHI provides practitioners with a decomposable, interpretable, and efficient faithfulness metric suitable for both offline evaluation and online monitoring of summarization systems.
Figures & tables
| Input Texts Source ( ) Reference ( ) Generated ( ) |
| Entity Extraction spaCy NER Relation Extraction SVO dep. parsing + Sentence-BERT matching Quantity Extraction Regex + spaCy NER Tolerance-aware matching |
| 5 Factors EF, PH, OF, NH, LF 6 Factors EF, PH, OF, NH, LH, LF 5 Factors QEF, QPH, QOF, QNH, QLF |
| Softmax normalization |
| EHI Entity faithfulness RHI ∗ Relation faithfulness QHI Quantity faithfulness |
| 4 Domains 200 articles = 800 source documents |
| 5 Models : BART, T5, Extractive, Corrupted, Reference |
| 4000 summaries |
| CHI metrics Baselines EHI, RHI ∗ , QHI, CHI ROUGE, BERTScore, AlignScore, SummaC |
| SummEval Human Judgments 100 art. 16 sys., 3 annotators, 4 dims (1–5) |
| Domain | Source | N | Avg. Words | Avg. Quant. |
|---|---|---|---|---|
| News | XSUM | 200 | 410 | 3.2 |
| Medical | PubMed | 200 | 1,796 | 8.7 |
| Legal | CNN/DM (filtered) | 200 | 821 | 5.1 |
| Financial | CNN/DM (filtered) | 200 | 899 | 11.4 |
| Correlation | Spearman | Kendall |
|---|---|---|
| RHI ∗ vs. Linear RHI | 0.968 | 0.878 |
| RHI ∗ vs. Human (system-level) | 0.506 | 0.333 |
| Linear RHI vs. Human (system-level) | 0.510 | 0.340 |
| Metric | System-Level | Summary-Level | ||
|---|---|---|---|---|
| QHI | 0.39 | 0.27 | 0.080 | 0.064 |
| ROUGE-L (quant. subset) | 0.32 | 0.21 | 0.095 | 0.072 |
| BERTScore (quant. subset) | — | — | 0.067 | 0.048 |
| AlignScore (quant. subset) | — | — | 0.12 | 0.09 |
| EHI | RHI ∗ | QHI | |
| EHI | 1.000 | 0.195 | 0.149 |
| RHI ∗ | 0.195 | 1.000 | 0.100 |
| QHI | 0.149 | 0.100 | 1.000 |
| Mean pairwise : 0.148 (all ) | |||
| Consist. | Coher. | Relev. | Avg. | Rank | |
| CHI | 0.25 | 0.54 | 0.70 | 0.66 | 1 |
| EHI | 0.15 | 0.49 | 0.62 | 0.58 | 2 |
| RHI ∗ | 0.51 | 0.38 | 0.50 | 0.56 | 3 |
| ROUGE-2 | 0.13 | 0.38 | 0.62 | 0.55 | 4 |
| ROUGE-1 | 0.16 | 0.36 | 0.61 | 0.53 | 5 |
| MiniCheck | 0.35 | 0.32 | 0.38 | 0.41 | 6 |
| Metric | System-Level | Summary-Level | Cost/doc | Interpretable | Decomposable | ||
|---|---|---|---|---|---|---|---|
| ROUGE-L | 0.39 | 0.27 | 0.139 | 0.109 | $0.00 | ✓ | — |
| BERTScore | — | — | 0.097 | 0.076 | $0.00 | — | — |
| AlignScore | — | — | 0.140 | 0.105 | $0.00 | — | — |
| SummaC | — | — | 0.134 | 0.098 | $0.00 | — | — |
| MiniCheck | 0.41 | 0.30 | 0.380 | 0.289 | $0.00 | — | — |
| Configuration | System | Summary | ||
|---|---|---|---|---|
| EHI only | 0.58 | 0.47 | 0.058 | 0.046 |
| RHI ∗ only | 0.56 | 0.45 | — | — |
| QHI only | 0.39 | 0.27 | 0.080 | 0.064 |
| EHI + RHI ∗ | 0.62 | 0.48 | — | — |
| EHI + QHI | 0.60 | 0.46 | 0.084 | 0.066 |
| Human Dimension | CHI | Best Baseline | Key Dim. | |
|---|---|---|---|---|
| Consistency | 0.25 | RHI ∗ (0.51) | 0.26 | RHI ∗ |
| Coherence | 0.54 ∗ | EHI (0.49) | +0.05 | EHI |
| Relevance | 0.70 ∗∗ | ROUGE-2 (0.62) | +0.08 | All |
| Average (all 4) | 0.66 ∗∗ | EHI (0.58) | +0.08 | All |
| Metric | Time/doc (s) | Cost/doc ($) | GPU Required |
|---|---|---|---|
| ROUGE-L | 0.01 | 0.000 | No |
| BERTScore | 0.50 | 0.000 | Yes (recommended) |
| AlignScore | 0.80 | 0.000 | Yes (recommended) |
| MiniCheck | 7.6 | 0.000 | No † |
| CHI (ours) | 4.3 | 0.000 | No † |
| GPT-4-judge | 3.0 | 0.150 | No (API) |
| Domain | % Entity | % Relation | % Quantity |
|---|---|---|---|
| News | 38.2 | 42.5 | 19.3 |
| Medical | 28.4 | 51.8 | 19.8 |
| Legal | 33.6 | 48.9 | 17.5 |
| Financial | 25.1 | 35.7 | 39.2 |
| Metric | Good Summary | Bad Summary | Gap | Correct Rank? |
|---|---|---|---|---|
| EHI (plain) | 0.272 | 0.298 | 0.026 | NO |
| EHI (UMLS) | 0.470 | 0.345 | +0.125 | YES |
| CHI (plain) | 0.299 | 0.339 | 0.040 | NO |
| CHI (UMLS-enhanced) | 0.353 | 0.357 | 0.004 |