Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging. Evaluating EA systems at scale requires expensive expert annotation, making systematic assessment across diverse domains practically infeasible. LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied. We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests. We identify anchor bias, a failure mode in which judges invert discrimination when the system's decision label is visible. Label exposure causally collapses judge discrimination (J-ROC-AUC 0.12-0.87), while a label-free protocol recovers near-ceiling capability on distinctive-name datasets (0.93-1.00) and significant recovery on biomedical pairs (0.93-0.95). Counterfactual experiments confirm causality (FSR 53-99%) and reveal a frontier model paradox: stronger judges exhibit greater label sensitivity, not less. A blinded two-annotator human evaluation (102 pairs, Cohen's kappa=0.902) confirms this mechanism directly. We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA. Code and data are available at https://github.com/vaibhavalakshmiravideshik/llm-as-a-judge-entity-alignment.
Figures & tables
Figure 1: LLM-as-judge evaluation for entity alignment, illustrated using D-W-15K (DBpedia–Wikidata). The two KG entities share the name “The Fast and the Furious” and the same director, but refer to different things: Entity A (DBpedia) is the 2001 film; Entity B (Wikidata) is the entire franchise series. The correct alignment label is Non-Match . Black path (LLM as judge): The EA system correctly predicts Non-Match ; the LLM judge scores this decision 3/10, implying it considers Non-Match unjustified—anchoring on the shared name and director rather than reasoning from evidence. Red path (LLM as predictor): When the entity pair is sent directly to the LLM without any system label, it predicts Match with 8/10 confidence—again anchored on surface name similarity.
Table 1: Metrics for EA system meta-evaluation (Experiment 2). Under Prompts I–II the LLM scores how well-justified the visible system decision is; J-ROC-AUC measures correctness discrimination. Under Prompt III the LLM acts as a predictor; no label is shown and Score Gap is the primary metric.
Figure 3: Frontier model paradox: FSR% (label sensitivity under Prompts I and II, LLM as judge) vs. J-ROC-AUC (alignment prediction capability under Prompt III, LLM as predictor). Points in the paradox zone (upper-left and upper-right) combine high label sensitivity with high or low discrimination, confirming that stronger judges are more susceptible to anchor bias yet more capable when labels are withheld. Shape = judge; color = dataset.
Dataset–Judge
System
MAS
FSR%
95% CI
LFR%
DW 4m
EasyEA
3.55
64
[55, 73]
17
DW 4o
EasyEA
4.66
71
[62, 80]
29
DW O4
EasyEA
6.07
99
[97, 100]
50
DW 4m
NeuSymEA
4.34
62
[52, 71]
41
DW 4o
NeuSymEA
6.39
81
[73, 88]
46
DW O4
NeuSymEA
8.21
95
[90, 99]
47
Table 2: Counterfactual label-flip results under Prompt I: MAS, FSR%, 95% CI, and LFR% ( n=100 per condition, 50 MATCH + 50 NON MATCH). FSR exceeds 50% in every condition. Label sensitivity scales monotonically with model capability; Claude Opus 4.7 reaches FSR = 99% on DW under EasyEA and 95% under NeuSymEA. 95% bootstrap CIs ( B=10,000 ) for FSR confirm that all lower bounds exceed chance (50%) except DY 4o-mini EasyEA ([43, 63]), where label sensitivity is weakest but FSR point estimate still exceeds chance. Compare to Prompt II amplification in Table 11 .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Judge
Prompt
t
J-Prec
J-Rec
J-F1
J-ROC-AUC
GPT-4o-mini
I
5
0.506
0.980
0.668
0.528
6
0.500
0.828
0.623
7
0.505
0.584
0.542
8
0.528
0.536
0.532
9
0.579
0.264
0.363
II
5
0.491
0.964
0.650
0.244
Appendix
Table 3: EasyEA Cheng et al. (2025) threshold sweep (DW): J-Prec, J-Rec, J-F1 at t∈{5,…,9} , J-ROC-AUC, n=500 .
Table 5: EasyEA Cheng et al. (2025) threshold sweep (DY): J-Prec, J-Rec, J-F1 at t∈{5,…,9} , J-ROC-AUC, n=500 .
Judge
Prompt
t
J-Prec
J-Rec
J-F1
J-ROC-AUC
GPT-4o-mini
I
5
0.499
0.948
0.654
0.477
6
0.474
0.756
0.582
7
0.474
0.520
0.496
8
0.473
0.456
0.464
9
0.539
0.304
0.389
II
5
0.493
0.928
0.644
0.341
Appendix
Table 6: EasyEA Cheng et al. (2025) threshold sweep (BIO): J-Prec, J-Rec, J-F1 at t∈{5,…,9} , J-ROC-AUC, n=500 .
Dataset
System
Judge
Prompt
t∗
J-Prec
J-Rec
J-F1
J-ROC-AUC
Score Gap
DW
NeuSymEA
GPT-4o-mini
I
5
0.508
0.972
0.668
0.355
− 0.900
GPT-4o-mini
II
5
0.494
0.972
0.655
0.145
− 1.072
GPT-4o-mini
III
6
0.972
0.988
0.980
0.981
+ 8.248
GPT-4o
I
5
0.496
0.983
0.659
0.251
− 1.207
GPT-4o
II
5
0.503
0.992
0.668
0.165
− 1.000
GPT-4o
III
5
0.976
0.988
0.982
0.995
+ 8.300
Appendix
Table 7: NeuSymEA Chen et al. (2026) and NeuSymEA-NoRefine best-threshold summary: J-Prec, J-Rec, J-F1, J-ROC-AUC, and Score Gap at t∗=argmaxtJ-F1(t) across DW and DY ( n=500 ; BIO excluded, NeuSymEA Chen et al. (2026) not evaluated on BIO). Discriminates: Yes = J-ROC-AUC >0.6 ; Invert = J-ROC-AUC <0.45 ; ∼ Chance otherwise.
Exp.
Cond.
Calls
In(M)
Out(M)
Cost
Avg
Exp. 1
DW-P1
12,000
21.4
9.8
$68.47
$0.0057
DW-P2
12,000
4.4
4.2
$21.92
$0.0018
DW-P3
12,000
17.0
9.7
$57.09
$0.0048
DY-P1
12,000
21.8
10.2
$70.70
$0.0059
DY-P2
12,000
4.3
3.9
$20.65
$0.0017
DY-P3
12,000
16.3
9.4
$57.53
$0.0048
Appendix
Table 8: API usage and cost by experiment, dataset, prompt, model.
Dataset
Judge
Prompt
MRI@1
MRI@2
MRI@3
DW
4o-mini
I
21.2
11.8
4.8
II
4.4
2.4
1.8
III
6.2
4.4
3.6
4o
I
5.0
2.6
0.4
II
4.6
2.8
2.4
III
3.4
2.0
1.0
Appendix
Table 9: MRI% for P1 Identity Anonymization at thresholds t∈{1,2,3} ( n=500 per condition). Bold values correspond to the t=2 threshold reported in the main paper. The domain asymmetry between BIO and general-domain datasets is preserved across all thresholds.
Data
Judge
P
PBI [CI]
MRI [CI]
DW
4o-mini
I
0.368 [0.222, 0.514]
11.8 [9.0, 14.8]
II
0.424 [0.344, 0.512]
2.4 [1.2, 3.8]
III
0.580 [0.466, 0.706]
4.4 [2.8, 6.4]
4o
I
− 0.050 [ − 0.138, 0.036]
2.6 [1.4, 4.0]
II
0.332 [0.254, 0.416]
2.8 [1.4, 4.4]
III
0.190 [0.118, 0.266]
2.0 [1.0, 3.4]
Appendix
Table 10: 95% bootstrap CIs ( B=10,000 ) for PBI and MRI under P1 ( n=500 ). Format: point [lower, upper]. PBI values in this table reflect the full resampling-based point estimates and differ from the raw point estimates in Table 14 due to estimation methodology; this table provides the authoritative PBI and MRI confidence intervals.
Dataset–Judge
System
MAS
FSR
LFR
DW 4m
EasyEA
6.92
85
45
DW 4o
EasyEA
7.68
94
50
DW O4
EasyEA
8.75
99
50
DW 4m
NeuSymEA
6.75
82
38
DW 4o
NeuSymEA
7.89
99
51
DW O4
NeuSymEA
9.64
99
50
Appendix
Table 11: Label-flip effects under Prompt II: MAS, FSR%, LFR% ( n=100 per condition). EasyEA Cheng et al. (2025) on DW, DY, BIO; NeuSymEA Chen et al. (2026) on DW and DY. Compare to Prompt I results in Table 2 .
Property
KG1 (MeSH)
KG2 (SNOMED CT)
Gold alignment pairs (sample)
15,000
Full alignment pairs (pre-sample)
46,823
Gold entities
13,219
14,866
Attribute triples
11,385,523
1,391,104
Relation triples
6,948,511
1,331,550
Entities in relation graph
2,456,953
379,282
Appendix
Table 12: BIO dataset statistics. Gold alignment size reflects the 15K stratified sample; ent_links_uri contains the full 46,823 pairs prior to sampling.
Data
Judge
Sys
P-I
P-II
P-III
DW
4o-mini
E
+ 0.14
− 0.79 ∗∗∗
+ 8.15 ∗∗∗
Emb
+ 0.23
− 0.76 ∗∗∗
+ 7.74 ∗∗∗
N
− 0.90 ∗∗∗
− 1.07 ∗∗∗
+ 8.25 ∗∗∗
NR
− 0.73 ∗∗∗
− 0.95 ∗∗∗
+ 8.25 ∗∗∗
4o
E
− 0.32
− 0.92 ∗∗∗
+ 8.09 ∗∗∗
Emb
− 0.40 ∗
− 0.80 ∗∗∗
+ 8.16 ∗∗∗
Appendix
Table 13: Score Gap ( Δ=sˉMATCH−sˉNON MATCH ) across two EA systems with ablations, three datasets, three judges, and three prompts ( n=500 per condition). ∗∗∗p<0.001 , ∗∗p<0.01 , ∗p<0.05 , n.s. p≥0.05 (Mann–Whitney U , two-sided). E=EasyEA, Emb=EasyEA-EmbOnly, N=NeuSymEA, NR=NeuSymEA-NoRefine. BIO: N and NR not evaluated ( ).
P1 — Identity Anon.
P2 — Supp. 30%
P3 — Supp. 50%
P6 — ID Removal
Dataset
Judge
Prompt
PBI
JSR
MRI%
PBI
JSR
MRI%
PBI
JSR
MRI%
PBI
JSR
MRI%
DW
4o-mini
I
+ 0.369
0.553
11.8
+ 0.439
0.516
11.2
+ 0.691
0.492
14.2
+ 0.679
0.549
13.2
II
+ 0.424
0.606
2.4
+ 0.180
0.606
0.2
+ 0.290
0.490
1.2
+ 0.288
0.567
1.2
III
+ 0.580
0.511
4.4
+ 0.174
0.453
0.4
+ 0.228
0.428
0.6
+ 0.266
0.412
0.8
4o
I
− 0.059
0.557
2.4
+ 0.313
0.502
1.8
+ 0.522
0.433
6.5
+ 0.398
0.434
5.5
II
+ 0.332
0.695
2.8
+ 0.298
0.530
1.0
+ 0.409
0.458
1.2
+ 0.268
0.536
0.8
Appendix
Table 14: Perturbation bias diagnostics under P1 (Identity Anonymization), P2 (Evidence Suppression 30%), P3 (Evidence Suppression 50%), and P6 (Identifier Removal) across all datasets, judges, and prompts ( n=500 per condition). PBI = mean score drop; JSR = Spearman rank correlation of scores before/after perturbation; MRI% = fraction of pairs with score collapse >2 pts. Bold = BIO conditions showing catastrophic collapse. P2 and P3 together establish a dose-response pattern for evidence suppression; P6 reveals unexpected identifier dependence in the biomedical domain. Full results for P4, P5 (negative control), and P7 are in Appendix G (Table 15 ).
P4 — Surface Transform.
P5 — Order Random. (control)
P7 — Struct.-to-Prose
Dataset
Judge
Prompt
PBI
JSR
MRI%
PBI
JSR
MRI%
PBI
JSR
MRI%
DW
4o-mini
I
+ 0.321
0.606
6.6
+ 0.000
0.578
4.8
− 0.140
0.544
5.2
II
+ 0.090
0.662
0.4
− 0.044
0.718
0.2
− 0.172
0.609
0.0
III
+ 0.112
0.497
0.4
+ 0.066
0.521
0.6
− 0.180
0.509
0.0
4o
I
+ 0.225
0.487
2.8
+ 0.058
0.540
0.2
− 0.030
0.515
0.6
II
+ 0.226
0.558
0.8
+ 0.030
0.660
0.6
− 0.174
0.607
0.4
Appendix
Table 15: Perturbation bias diagnostics for P4 (Surface Transformation), P5 (Order Randomization, negative control), and P7 (Structured-to-Prose) across all datasets, judges, and prompts ( n=500 per condition). PBI = mean score drop; JSR = Spearman rank correlation of scores before/after perturbation; MRI% = fraction of pairs with score collapse >2 pts. P5 MRI% <5 % across all conditions confirms the negative control is clean. P4 and P7 show modest general-domain effects and near-zero BIO MRI, confirming that the catastrophic BIO collapse in Table 14 is specific to identity anonymization.
DS
Judge
sˉ0 -I
sˉ0 -II
PDV-I
PDV-II
Ratio
ρ
DW
4o-mini
7.81
9.23
1.64
1.21
0.74
+ 0.017
4o
8.68
9.29
1.26
1.01
0.81
+ 0.314 ∗∗∗
Claude
8.88
9.81
1.32
0.71
0.54
+ 0.173 ∗∗∗
DY
4o-mini
7.60
9.21
1.70
0.66
0.39
+ 0.004
4o
7.45
8.94
1.49
0.77
0.52
+ 0.232 ∗∗∗
Claude
7.59
9.90
1.54
0.32
0.21
+ 0.093 ∗
Appendix
Table 16: Baseline score inflation and scoring variance under Prompts I and II ( n=500 per condition). sˉ0 = mean baseline score; PDV = prompt discriminability variance; PDV ratio =PDVII/PDVI ; ρ = Spearman correlation between prompt and perturbation sensitivity. ∗∗∗p<0.001 , ∗p<0.05 , n.s. p≥0.05 .
Inter-annotator agreement
ρ
κ
MAE
Overall (all prompts, n=306 )
0.967
0.902
0.38
Match pairs
—
0.731
—
Non-Match pairs
—
0.891
—
Gold-label accuracy (Prompt III)
Match scored ≥6
100.0%
Non-Match scored <6
82.7%
Appendix
Table 17: Human evaluation summary (Experiment 4; n=102 pairs, 306 scored instances per annotator). Top: blinded inter-annotator agreement, overall and by true label. Bottom: correlation between the human consensus (mean of both annotators) and each LLM judge under Prompt I, overall and by true label. The MATCH/NON-MATCH split shows the anchor-bias signature directly: agreement with humans collapses on MATCH pairs, where the visible label most often disagrees with careful evidence-based judgment, and stays high on NON-MATCH pairs.