Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges
Authors: Donghao Huang, Jinling Pei, Zhaoxia Wang
Organizations: Research and Development, Mastercard, Arlington, VA, USA · School of Computing and Information Systems, Singapore Management University, Singapore
Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels contaminate pseudo-labeled data, while conservative labeling reduces coverage. We investigate whether agreement between two local large language model judges improves pseudo-label reliability. A label is retained only when the judges agree, with separate ordered thresholds for selections and abstentions that guarantee disjoint positive and negative label sets. Retrospective replay on 2,000 expert-annotated queries shows that higher selection thresholds can improve positive-label purity, whereas higher abstention thresholds increase false no-match labels. At thresholds (0.86, 0.80), Muse Glimmer 30B and Gemma 4 31B jointly label 1,633 queries (81.7% coverage) at 96.88% purity; positive and negative purities are 99.47% and 93.38%. This exceeds either constituent model at the same thresholds by more than two percentage points, with lower coverage. A split-half check finds only 0.14 percentage points of threshold-selection optimism. A symmetric threshold of 0.86 adds 40 erroneous no-match labels, while 46 false abstentions persist even with no confidence threshold. Across five matched within-model comparisons, higher reasoning effort yields no clear F0.5 gain and increases median latency by 1.8-5.0 times. These results motivate separate thresholding and auditing for positive and negative pseudo-labels. The study establishes label purity, not student utility; fresh-data curation and student fine-tuning remain necessary to demonstrate downstream value.
Figures & tables
Judge
Configuration
F0.5
P
Rm
N-FMR
s/row
GPT OSS 120B
deployed, medium
94.83
95.94
90.64
6.03
–
Muse Glimmer 30B
zero-shot, low, v1
96.76
97.89
92.48
2.16
35.2
Gemma 4 31B
k=20 , off, v2
95.27
95.69
93.63
5.75
22.2
GPT OSS 120B
low, v1 (local)
94.95
95.89
91.33
5.60
16.9
Claude Sonnet 4.5
low, v1 (API)
97.41
97.81
95.86
2.59
14.2
TABLE I: Judges on the frozen 2,000-row benchmark ( τ=.70 ). Quality metrics are percentages. Low/medium are provider-specific effort settings; off denotes disabled thinking. Time is median serial model-call latency (local workstation or Sonnet API), not production throughput or cost.
Curator
Positives
Purity
Negatives
Purity
Total
Purity
Coverage
Review
Muse Glimmer 30B
1,077
99.26
773
88.62
1,850
94.81
92.5
150
Gemma 4 31B
1,175
97.19
725
90.48
1,900
94.63
95.0
100
GPT OSS 120B
1,212
96.78
765
86.41
1,977
92.77
98.8
23
Claude Sonnet 4.5
1,257
98.25
726
93.66
1,983
96.57
99.2
17
Muse × Gemma
938
99.47
695
93.38
1,633
96.88
81.7
367
TABLE II: Curated labels at τsel=0.86 , τabs=0.80 . Purity and coverage are percentages; review is the remainder of the 2,000 rows. The policies select different subsets, so this comparison does not hold coverage fixed.
Policy
Err. +
Err. −
Purity
Coverage
Symmetric .70
12
46
96.77
89.7
Symmetric .80
11
46
96.82
89.7
Symmetric .86
5
86
94.64
84.8
Asymmetric
5
46
96.88
81.7
Positive only
5
0
99.47
46.9
TABLE III: Gate ablation with fixed outputs. Asymmetric uses (.86,.80) ; positive-only uses τsel=.86 . Err. + and Err. − count incorrect labels; purity and coverage are percentages.
τsel\τabs
0.70
0.75
0.80
0.86
0.80
96.82
96.82
96.82
—
0.86
96.87
96.88
96.88
94.64
0.90
96.83
96.84
96.84
94.57
0.95
95.95
95.95
95.96
92.90
TABLE IV: Total label purity (%) over the gate grid. The excluded cell violates the protocol’s disjointness guarantee and overlaps on 19 rows. Bold marks the illustrative operating point; polarity-specific quantities are separable, but total purity also depends on label counts.
Policy
τsel
fi
Total
Pos.
Neg.
Cov.
Inherited (full)
0.86
0.80
96.88
99.47
93.38
81.7
Objective optimum (full)
0.86
0.81
96.88
99.47
93.40
81.8
Objective-selected (held-out)
0.60–0.86
0.60–0.81
96.79
99.14
93.39
85.5
Inherited (held-out)
0.86
0.80
96.88
99.47
93.39
81.7
TABLE V: Held-out check of the operating point. The stated objective maximizes total purity subject to coverage ≥ 80% over the valid grid. Held-out rows average 400 evaluations from 200 stratified random half-splits (seed 42): thresholds are selected on one half and scored on the other. This quantifies threshold-selection optimism only; model and prompt selection also reused this benchmark. Purity and coverage are percentages.
Recent large language models (LLMs) achieve strong performance on entity matching without requiring task-specific training data. However, applying these models to large sets of candidate pairs remains slow and costly. In contrast, entity matchers using traditional machine learning methods or small language models (SLMs), such as RoBERTa, offer much faster inference but require task-specific training data. This paper investigates whether the need to provide task-specific training data can be avoided by using knowledge-distillation workflows, in which an LLM serves as a teacher model to label training pairs that are subsequently used to train a smaller student model. We investigate knowledge distillation for entity matching along the following dimensions: pair-selection strategy, teacher model, label post-processing method, and student model. We evaluate the workflows using the Abt-Buy, Walmart-Amazon, WDC Products, DBLP-ACM, and DBLP-Scholar benchmarks, and compare the performance of student models trained with machine-labeled data to the performance of the same models trained using the benchmark training sets. Our experiments show that student models trained using the machine-labeled sets perform approximately on par with models trained on the benchmark training sets, with the remaining differences in both directions staying below two F1 points. Using GPT-5.2 to label the training sets for all five benchmarks costs US$28.31 to US$40.88, whereas manually labeling the same training sets is estimated to require 470 hours of work. At inference time, Ditto is 41.5 to 534 times faster than directly using an LLM to perform the matching tasks. These results indicate that current LLMs, when combined with a suitable pair-selection method, can substantially reduce or even eliminate the manual effort required to label use case-specific training data for entity matching.
Aaron Steiner, Christian Bizer
Data and Web Science Group University of Mannheim Mannheim, Germany
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignment with human raters -- a property that itself depends on costly human annotations. In this work, we develop a method (Metric Match) for estimating correlation-based reliability metrics of LLM judges from limited annotations. Metric Match selects a subset of samples for human annotation such that the subset matches the population reliability metric with respect to acquired synthetic labels. We empirically show that Metric Match achieves a win-rate of 0.838 against random subset selection across four different correlation metrics and 15 datasets, with an 18.7% decrease in average estimation error and reduces annotation needs by 32.5%. We provide a cost model and highlight a medical case study where our method saves $1,041.67 compared to random selection for expert annotation. Further, we shift our task from reliability estimation to reliability classification of whether a given judge is above a deployment threshold, outperforming random selection with Metric Match. All project code is publicly available, and we additionally provide an installable package for ease of use.
Alyssa Unell, Natalie Dullerud, Naomi Boneh +4
Department of Computer Science Stanford University · Department of Medicine Stanford University
Multi-judge evaluation is increasingly used to assess LLMs and reward models, and the prevailing heuristic is to curate: keep the most accurate judges and discard weaker ones. We show that this heuristic can reverse when the target is not point accuracy, but calibrated probabilistic evaluation from a labeled calibration set. Holding the aggregation and calibration procedures fixed, we compare accuracy-ranked top-k judge selection with using the full judge panel. Across four labeled pairwise-evaluation benchmarks spanning LLM-as-judge and reward-model settings, the calibrated full panel consistently outperforms accuracy-based selection. On RewardBench2, retaining all judges achieves negative log-likelihood (NLL) of 0.006 versus 0.013 under top-5 selection, halving the calibration error. This advantage persists after judge-family deduplication and against stronger same-pipeline subset search. We explain this reversal with oracle analyses showing that the optimal calibrated risk under proper scoring rules cannot increase when additional judge signals are made available, and that even below-chance judges can be useful when their biases are learnable and their signals are non-redundant. The resulting operating principle is simple: in multi-judge evaluation with labeled calibration data, do not discard weak judges by accuracy alone; keep them when they are parseable, non-redundant, and calibratable.