cs.AIOct 7, 2026

Reliability of LLM Judges for Evaluating Entity Alignment

Authors: Vaibhava Lakshmi Ravideshik, Mayank Kejriwal

Organizations: University of Michigan, Ann Arbor · Information Sciences Institute, University of Southern California

Abstract

Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging. Evaluating EA systems at scale requires expensive expert annotation, making systematic assessment across diverse domains practically infeasible. LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied. We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests. We identify anchor bias, a failure mode in which judges invert discrimination when the system's decision label is visible. Label exposure causally collapses judge discrimination (J-ROC-AUC 0.12-0.87), while a label-free protocol recovers near-ceiling capability on distinctive-name datasets (0.93-1.00) and significant recovery on biomedical pairs (0.93-0.95). Counterfactual experiments confirm causality (FSR 53-99%) and reveal a frontier model paradox: stronger judges exhibit greater label sensitivity, not less. A blinded two-annotator human evaluation (102 pairs, Cohen's kappa=0.902) confirms this mechanism directly. We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA. Code and data are available at https://github.com/vaibhavalakshmiravideshik/llm-as-a-judge-entity-alignment.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Predicate Importance Estimation and Decoupled Rationale-Score Distillation for Entity Alignment

    Jun 22, 2026Keunha Kim, Yoonjin Jang, Hyeon-gu Lee +2KG EmbeddingKnowledge Graphs

  2. HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment

    May 27, 2026Yoonjin Jang, Junwoo Kim, Youngjoong KoHard Negative Mining