Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging. Evaluating EA systems at scale requires expensive expert annotation, making systematic assessment across diverse domains practically infeasible. LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied. We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests. We identify anchor bias, a failure mode in which judges invert discrimination when the system's decision label is visible. Label exposure causally collapses judge discrimination (J-ROC-AUC 0.12-0.87), while a label-free protocol recovers near-ceiling capability on distinctive-name datasets (0.93-1.00) and significant recovery on biomedical pairs (0.93-0.95). Counterfactual experiments confirm causality (FSR 53-99%) and reveal a frontier model paradox: stronger judges exhibit greater label sensitivity, not less. A blinded two-annotator human evaluation (102 pairs, Cohen's kappa=0.902) confirms this mechanism directly. We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA. Code and data are available at https://github.com/vaibhavalakshmiravideshik/llm-as-a-judge-entity-alignment.
Figures & tables
Figure 1: LLM-as-judge evaluation for entity alignment, illustrated using D-W-15K (DBpedia–Wikidata). The two KG entities share the name “The Fast and the Furious” and the same director, but refer to different things: Entity A (DBpedia) is the 2001 film; Entity B (Wikidata) is the entire franchise series. The correct alignment label is Non-Match . Black path (LLM as judge): The EA system correctly predicts Non-Match ; the LLM judge scores this decision 3/10, implying it considers Non-Match unjustified—anchoring on the shared name and director rather than reasoning from evidence. Red path (LLM as predictor): When the entity pair is sent directly to the LLM without any system label, it predicts Match with 8/10 confidence—again anchored on surface name similarity.
Table 1: Metrics for EA system meta-evaluation (Experiment 2). Under Prompts I–II the LLM scores how well-justified the visible system decision is; J-ROC-AUC measures correctness discrimination. Under Prompt III the LLM acts as a predictor; no label is shown and Score Gap is the primary metric.
Figure 3: Frontier model paradox: FSR% (label sensitivity under Prompts I and II, LLM as judge) vs. J-ROC-AUC (alignment prediction capability under Prompt III, LLM as predictor). Points in the paradox zone (upper-left and upper-right) combine high label sensitivity with high or low discrimination, confirming that stronger judges are more susceptible to anchor bias yet more capable when labels are withheld. Shape = judge; color = dataset.
Dataset–Judge
System
MAS
FSR%
95% CI
LFR%
DW 4m
EasyEA
3.55
64
[55, 73]
17
DW 4o
EasyEA
4.66
71
[62, 80]
29
DW O4
EasyEA
6.07
99
[97, 100]
50
DW 4m
NeuSymEA
4.34
62
[52, 71]
41
DW 4o
NeuSymEA
6.39
81
[73, 88]
46
DW O4
NeuSymEA
8.21
95
[90, 99]
47
Table 2: Counterfactual label-flip results under Prompt I: MAS, FSR%, 95% CI, and LFR% ( n=100 per condition, 50 MATCH + 50 NON MATCH). FSR exceeds 50% in every condition. Label sensitivity scales monotonically with model capability; Claude Opus 4.7 reaches FSR = 99% on DW under EasyEA and 95% under NeuSymEA. 95% bootstrap CIs ( B=10,000 ) for FSR confirm that all lower bounds exceed chance (50%) except DY 4o-mini EasyEA ([43, 63]), where label sensitivity is weakest but FSR point estimate still exceeds chance. Compare to Prompt II amplification in Table 11 .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Judge
Prompt
t
J-Prec
J-Rec
J-F1
J-ROC-AUC
GPT-4o-mini
I
5
0.506
0.980
0.668
0.528
6
0.500
0.828
0.623
7
0.505
0.584
0.542
8
0.528
0.536
0.532
9
0.579
0.264
0.363
II
5
0.491
0.964
0.650
0.244
Appendix
Table 3: EasyEA Cheng et al. (2025) threshold sweep (DW): J-Prec, J-Rec, J-F1 at t∈{5,…,9} , J-ROC-AUC, n=500 .
Table 5: EasyEA Cheng et al. (2025) threshold sweep (DY): J-Prec, J-Rec, J-F1 at t∈{5,…,9} , J-ROC-AUC, n=500 .
Judge
Prompt
t
J-Prec
J-Rec
J-F1
J-ROC-AUC
GPT-4o-mini
I
5
0.499
0.948
0.654
0.477
6
0.474
0.756
0.582
7
0.474
0.520
0.496
8
0.473
0.456
0.464
9
0.539
0.304
0.389
II
5
0.493
0.928
0.644
0.341
Appendix
Table 6: EasyEA Cheng et al. (2025) threshold sweep (BIO): J-Prec, J-Rec, J-F1 at t∈{5,…,9} , J-ROC-AUC, n=500 .
Dataset
System
Judge
Prompt
t∗
J-Prec
J-Rec
J-F1
J-ROC-AUC
Score Gap
DW
NeuSymEA
GPT-4o-mini
I
5
0.508
0.972
0.668
0.355
− 0.900
GPT-4o-mini
II
5
0.494
0.972
0.655
0.145
− 1.072
GPT-4o-mini
III
6
0.972
0.988
0.980
0.981
+ 8.248
GPT-4o
I
5
0.496
0.983
0.659
0.251
− 1.207
GPT-4o
II
5
0.503
0.992
0.668
0.165
− 1.000
GPT-4o
III
5
0.976
0.988
0.982
0.995
+ 8.300
Appendix
Table 7: NeuSymEA Chen et al. (2026) and NeuSymEA-NoRefine best-threshold summary: J-Prec, J-Rec, J-F1, J-ROC-AUC, and Score Gap at t∗=argmaxtJ-F1(t) across DW and DY ( n=500 ; BIO excluded, NeuSymEA Chen et al. (2026) not evaluated on BIO). Discriminates: Yes = J-ROC-AUC >0.6 ; Invert = J-ROC-AUC <0.45 ; ∼ Chance otherwise.
Exp.
Cond.
Calls
In(M)
Out(M)
Cost
Avg
Exp. 1
DW-P1
12,000
21.4
9.8
$68.47
$0.0057
DW-P2
12,000
4.4
4.2
$21.92
$0.0018
DW-P3
12,000
17.0
9.7
$57.09
$0.0048
DY-P1
12,000
21.8
10.2
$70.70
$0.0059
DY-P2
12,000
4.3
3.9
$20.65
$0.0017
DY-P3
12,000
16.3
9.4
$57.53
$0.0048
Appendix
Table 8: API usage and cost by experiment, dataset, prompt, model.
Dataset
Judge
Prompt
MRI@1
MRI@2
MRI@3
DW
4o-mini
I
21.2
11.8
4.8
II
4.4
2.4
1.8
III
6.2
4.4
3.6
4o
I
5.0
2.6
0.4
II
4.6
2.8
2.4
III
3.4
2.0
1.0
Appendix
Table 9: MRI% for P1 Identity Anonymization at thresholds t∈{1,2,3} ( n=500 per condition). Bold values correspond to the t=2 threshold reported in the main paper. The domain asymmetry between BIO and general-domain datasets is preserved across all thresholds.
Data
Judge
P
PBI [CI]
MRI [CI]
DW
4o-mini
I
0.368 [0.222, 0.514]
11.8 [9.0, 14.8]
II
0.424 [0.344, 0.512]
2.4 [1.2, 3.8]
III
0.580 [0.466, 0.706]
4.4 [2.8, 6.4]
4o
I
− 0.050 [ − 0.138, 0.036]
2.6 [1.4, 4.0]
II
0.332 [0.254, 0.416]
2.8 [1.4, 4.4]
III
0.190 [0.118, 0.266]
2.0 [1.0, 3.4]
Appendix
Table 10: 95% bootstrap CIs ( B=10,000 ) for PBI and MRI under P1 ( n=500 ). Format: point [lower, upper]. PBI values in this table reflect the full resampling-based point estimates and differ from the raw point estimates in Table 14 due to estimation methodology; this table provides the authoritative PBI and MRI confidence intervals.
Dataset–Judge
System
MAS
FSR
LFR
DW 4m
EasyEA
6.92
85
45
DW 4o
EasyEA
7.68
94
50
DW O4
EasyEA
8.75
99
50
DW 4m
NeuSymEA
6.75
82
38
DW 4o
NeuSymEA
7.89
99
51
DW O4
NeuSymEA
9.64
99
50
Appendix
Table 11: Label-flip effects under Prompt II: MAS, FSR%, LFR% ( n=100 per condition). EasyEA Cheng et al. (2025) on DW, DY, BIO; NeuSymEA Chen et al. (2026) on DW and DY. Compare to Prompt I results in Table 2 .
Property
KG1 (MeSH)
KG2 (SNOMED CT)
Gold alignment pairs (sample)
15,000
Full alignment pairs (pre-sample)
46,823
Gold entities
13,219
14,866
Attribute triples
11,385,523
1,391,104
Relation triples
6,948,511
1,331,550
Entities in relation graph
2,456,953
379,282
Appendix
Table 12: BIO dataset statistics. Gold alignment size reflects the 15K stratified sample; ent_links_uri contains the full 46,823 pairs prior to sampling.
Data
Judge
Sys
P-I
P-II
P-III
DW
4o-mini
E
+ 0.14
− 0.79 ∗∗∗
+ 8.15 ∗∗∗
Emb
+ 0.23
− 0.76 ∗∗∗
+ 7.74 ∗∗∗
N
− 0.90 ∗∗∗
− 1.07 ∗∗∗
+ 8.25 ∗∗∗
NR
− 0.73 ∗∗∗
− 0.95 ∗∗∗
+ 8.25 ∗∗∗
4o
E
− 0.32
− 0.92 ∗∗∗
+ 8.09 ∗∗∗
Emb
− 0.40 ∗
− 0.80 ∗∗∗
+ 8.16 ∗∗∗
Appendix
Table 13: Score Gap ( Δ=sˉMATCH−sˉNON MATCH ) across two EA systems with ablations, three datasets, three judges, and three prompts ( n=500 per condition). ∗∗∗p<0.001 , ∗∗p<0.01 , ∗p<0.05 , n.s. p≥0.05 (Mann–Whitney U , two-sided). E=EasyEA, Emb=EasyEA-EmbOnly, N=NeuSymEA, NR=NeuSymEA-NoRefine. BIO: N and NR not evaluated ( ).
P1 — Identity Anon.
P2 — Supp. 30%
P3 — Supp. 50%
P6 — ID Removal
Dataset
Judge
Prompt
PBI
JSR
MRI%
PBI
JSR
MRI%
PBI
JSR
MRI%
PBI
JSR
MRI%
DW
4o-mini
I
+ 0.369
0.553
11.8
+ 0.439
0.516
11.2
+ 0.691
0.492
14.2
+ 0.679
0.549
13.2
II
+ 0.424
0.606
2.4
+ 0.180
0.606
0.2
+ 0.290
0.490
1.2
+ 0.288
0.567
1.2
III
+ 0.580
0.511
4.4
+ 0.174
0.453
0.4
+ 0.228
0.428
0.6
+ 0.266
0.412
0.8
4o
I
− 0.059
0.557
2.4
+ 0.313
0.502
1.8
+ 0.522
0.433
6.5
+ 0.398
0.434
5.5
II
+ 0.332
0.695
2.8
+ 0.298
0.530
1.0
+ 0.409
0.458
1.2
+ 0.268
0.536
0.8
Appendix
Table 14: Perturbation bias diagnostics under P1 (Identity Anonymization), P2 (Evidence Suppression 30%), P3 (Evidence Suppression 50%), and P6 (Identifier Removal) across all datasets, judges, and prompts ( n=500 per condition). PBI = mean score drop; JSR = Spearman rank correlation of scores before/after perturbation; MRI% = fraction of pairs with score collapse >2 pts. Bold = BIO conditions showing catastrophic collapse. P2 and P3 together establish a dose-response pattern for evidence suppression; P6 reveals unexpected identifier dependence in the biomedical domain. Full results for P4, P5 (negative control), and P7 are in Appendix G (Table 15 ).
P4 — Surface Transform.
P5 — Order Random. (control)
P7 — Struct.-to-Prose
Dataset
Judge
Prompt
PBI
JSR
MRI%
PBI
JSR
MRI%
PBI
JSR
MRI%
DW
4o-mini
I
+ 0.321
0.606
6.6
+ 0.000
0.578
4.8
− 0.140
0.544
5.2
II
+ 0.090
0.662
0.4
− 0.044
0.718
0.2
− 0.172
0.609
0.0
III
+ 0.112
0.497
0.4
+ 0.066
0.521
0.6
− 0.180
0.509
0.0
4o
I
+ 0.225
0.487
2.8
+ 0.058
0.540
0.2
− 0.030
0.515
0.6
II
+ 0.226
0.558
0.8
+ 0.030
0.660
0.6
− 0.174
0.607
0.4
Appendix
Table 15: Perturbation bias diagnostics for P4 (Surface Transformation), P5 (Order Randomization, negative control), and P7 (Structured-to-Prose) across all datasets, judges, and prompts ( n=500 per condition). PBI = mean score drop; JSR = Spearman rank correlation of scores before/after perturbation; MRI% = fraction of pairs with score collapse >2 pts. P5 MRI% <5 % across all conditions confirms the negative control is clean. P4 and P7 show modest general-domain effects and near-zero BIO MRI, confirming that the catastrophic BIO collapse in Table 14 is specific to identity anonymization.
DS
Judge
sˉ0 -I
sˉ0 -II
PDV-I
PDV-II
Ratio
ρ
DW
4o-mini
7.81
9.23
1.64
1.21
0.74
+ 0.017
4o
8.68
9.29
1.26
1.01
0.81
+ 0.314 ∗∗∗
Claude
8.88
9.81
1.32
0.71
0.54
+ 0.173 ∗∗∗
DY
4o-mini
7.60
9.21
1.70
0.66
0.39
+ 0.004
4o
7.45
8.94
1.49
0.77
0.52
+ 0.232 ∗∗∗
Claude
7.59
9.90
1.54
0.32
0.21
+ 0.093 ∗
Appendix
Table 16: Baseline score inflation and scoring variance under Prompts I and II ( n=500 per condition). sˉ0 = mean baseline score; PDV = prompt discriminability variance; PDV ratio =PDVII/PDVI ; ρ = Spearman correlation between prompt and perturbation sensitivity. ∗∗∗p<0.001 , ∗p<0.05 , n.s. p≥0.05 .
Inter-annotator agreement
ρ
κ
MAE
Overall (all prompts, n=306 )
0.967
0.902
0.38
Match pairs
—
0.731
—
Non-Match pairs
—
0.891
—
Gold-label accuracy (Prompt III)
Match scored ≥6
100.0%
Non-Match scored <6
82.7%
Appendix
Table 17: Human evaluation summary (Experiment 4; n=102 pairs, 306 scored instances per annotator). Top: blinded inter-annotator agreement, overall and by true label. Bottom: correlation between the human consensus (mean of both annotators) and each LLM judge under Prompt I, overall and by true label. The MATCH/NON-MATCH split shows the anchor-bias signature directly: agreement with humans collapses on MATCH pairs, where the visible label most often disagrees with careful evidence-based judgment, and stays high on NON-MATCH pairs.
Knowledge graphs (KGs) are increasingly used as structured context for Large Language Models (LLMs), but industrial KG-RAG systems often need to integrate public and domain-specific KGs constructed from heterogeneous databases. This integration relies on Entity Alignment (EA), where lexical matching alone is insufficient under predicate-name variation and incomplete local neighborhoods. We address EA for KG integration by constructing a pairwise EA dataset and proposing two complementary modules: Predicate Importance Estimation (PIE) and Decoupled Rationale-Score Distillation (DRSD). PIE is a compact embedding-based approach that removes the subject information from each 1-hop triple, encodes the resulting subjectless triples, and aggregates them with learnable predicate-importance weights to build predicate-aware entity embeddings. DRSD trains a distilled small language model (SLM) with pseudo-answers produced by a teacher LLM through distinct prompts. By converting binary EA labels into text-based supervision and decoupling confidence-score estimation from label-consistent rationales, DRSD enables the SLM to learn task-specific reasoning while retaining a less label-biased confidence signal. Experiments show that PIE and DRSD improve EA classification. Moreover, because DRSD decouples confidence-score estimation from the decision, a discrepancy between the two flags an uncertain prediction for human review, thereby enabling a practical discrepancy between automatic acceptance and human-in-the-loop verification.
Entity Alignment (EA) is essential for knowledge graph (KG) fusion, but existing benchmarks often allow models to exploit name overlap rather than relational structure. This makes it difficult to evaluate whether models can reject same-name entities that refer to different real-world objects. Our primary contribution is a same-name hard-negative augmentation strategy that simultaneously yields quality-controlled evaluation benchmarks (DW-HN29K, DY-HN27K) and augmented training corpora (DW-Train, DY-Train), by mining same-name but distinct entity pairs from KG name-collision groups. We further introduce HELEA, a two-stage framework integrating (i) entity encoder retrieval trained on hard-negative-augmented training corpora with 1-hop KG context, and (ii) LLM-based reranking without additional training. Experiments show that name-dependent baselines collapse to near-random performance on our hard-negative benchmarks, while HELEA achieves F1 0.967 on DW-HN29K while maintaining Hit@1 0.993 on standard DW-15K.
Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-world object. Conventional EA methods mainly exploit explicit graph structures and textual fields, which often provide insufficient semantic understanding to recognize the same entity under heterogeneous descriptions and distinguish it from semantically similar entities. Although large language models (LLMs) offer deeper entity understanding, existing LLM-based EA methods largely use this capability for auxiliary generation or candidate-conditioned decisions. Consequently, such understanding is not distilled into a stable and directly comparable identity space, leaving alignment tied to specific KG pairs or candidate sets and requiring repeated processing as the matching context changes. To address these limitations, we propose IRIS (Identity Representations from Internal States), a training-free framework that constructs for each entity an iris-like signature encoding its distinctive and stable identity characteristics. IRIS derives these signatures by eliciting identity-oriented contextual representations from a frozen LLM, thereby forming a shared space in which each entity is encoded once and can be aligned across different KGs through direct similarity comparison, without pair-dependent representation construction or candidate-wise LLM inference. Across four established EA benchmarks and two frozen LLM backbones, the best IRIS variants achieve Hits@1 scores of 100.00, 99.38, 98.31, and 97.99 on D-Y-15K V2, DBP-WIKI, ICEWS-WIKI, and ICEWS-YAGO, respectively.