cs.AIOct 4, 2026

Better Retrieval, Limited Clustering Gains: A Controlled Study of Multilingual Company Entity Resolution

Authors: Yijiashun Qi, Yuxuan Li, Hanzhe Guo

Organizations: University of Michigan, Ann Arbor, USA · University of Pennsylvania, Philadelphia, PA, USA

Abstract

Improved name retrieval may have little effect on company clusters when the pair classifier remains unchanged. We examine this dependency by adapting multilingual E5 encoders under fixed candidate budgets and downstream decision rules. Random-negative and hard-negative training use identical positive schedules. Checkpoints are selected before collecting a new GLEIF sample of 3,633 names, 2,880 source identities and 882 silver-positive pairs. At 72,660 candidate edges, adaptation with a multi-view selector increases direct pair recall from 53.74% to 76.98%. The primary matcher adds only seven correct and two incorrect co-cluster pairs: cluster recall rises from 32.54% to 33.33%, while precision falls from 95.99% to 95.45%. Of 208 newly retrieved silver-positive pairs, 202 fall below its decision threshold. Random-negative and hard-negative training produce identical final partitions. An AI-assisted, single-reviewer audit of 137 pairs supports the observed pattern, although its predominantly LEI-derived evidence does not establish independent gold labels. The results locate the immediate loss of retrieval gains at the existing confirmation stage and show why encoder evaluation must also measure final cluster quality.

Explore similar work

CardsList
  1. HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment

    May 27, 2026Yoonjin Jang, Junwoo Kim, Youngjoong KoKnowledge GraphsReranking

  2. Entity Resolution in Practice: Lessons from a Self-Serve Pipeline

    Jul 28, 2026Kaushik Pavani, Ganga Aluri, Pravin Jadhav +2Data MiningFalse-Negative Rate