Billiger.de Products: A Bilingual Entity Matching Benchmark
Authors: Aaron Steiner, Ksenia Elagin, Ralph Peeters, Johannes Knopp, Christian Bizer
Organizations: University of Mannheim, Data and Web Science Group, B6, 26, 68159 Mannheim, Germany · solute GmbH, Zeppelinstraße 15, 76185 Karlsruhe, Germany
Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces Billiger.de Products, a bilingual German and English entity matching benchmark covering thirteen consumer product categories, including difficult-to-handle categories such as clothing and furniture. The benchmark data originates from the German price comparison platform billiger.de. Following the design of WDC Products, the benchmark offers multiple variants that differ in the fraction of corner cases, the size of the development set, and the fraction of entities unseen during training. An aligned English translation of every offer keeps all pairs, splits, and labels fixed, while cross-language test sets combine German and English records within individual pairs. We validate the benchmark using six supervised matchers and zero-shot GPT-5.2 on both language versions and the cross-language test sets. The validation shows the difficulty of the benchmark. The comparison of the results on the English version of the benchmark to the results on the German version shows that most matchers score on average higher on the English version. The difference is largest for RoBERTa and HierGAT, while the zero-shot LLM runs are largely insensitive to the language. Comparing the F1 scores achieved by PLM-based matchers on the English version of Billiger.de Products with their performance on existing English-language benchmarks, such as WDC Products and Abt-Buy, shows that Billiger.de Products is more difficult than these benchmarks.
Figures & tables
Figure 2: Construction pipeline of the benchmark, from clustering to the aligned German and English releases.
Training
Validation
Test
CC
Size
Pos.
Neg.
Total
Pos.
Neg.
Total
Unseen
Pos.
Neg.
Total
20 %
Small
485
1,979
2,464
491
1,993
2,484
0 %
407
4,038
4,445
Medium
1,457
4,456
5,913
491
2,989
3,480
50 %
347
4,082
4,429
Large
13,100
14,012
27,112
491
3,984
4,475
100 %
360
4,089
4,449
50 %
Small
482
1,979
2,461
479
1,986
2,465
0 %
363
4,049
4,412
Medium
1,464
4,441
5,905
479
2,977
3,456
50 %
326
4,092
4,418
Table 1: Pair composition per corner-case ratio (CC), identical for both languages. Validation reports the Seen splits used for model selection. Training and validation vary by development set size, while test sets vary by unseen-product ratio.
Category
Train (%)
Test (%)
Category
Train (%)
Test (%)
Clothing & Accessories
20.85
28.56
Cosmetics & Drugstore
3.25
4.51
Furniture & Living
21.10
26.42
Health & Care
2.24
2.46
Electronics & Computers
20.53
12.51
Office Supplies
0.94
0.69
Tools & DIY
11.34
8.10
Food & Beverages
1.00
0.68
Sports & Leisure
6.23
6.22
Books, Films & Music
0.49
0.51
Toys & Baby
6.24
4.92
Pet Supplies
0.23
0.19
Table 2: Share of product offers per category in the training and test splits in percent, averaged over all variants.
Supervised
GPT-5.2 zero-shot
CC
Size
Test
WordCooc
Magellan
RoBERTa
R-SupCon
HierGAT
Ditto
Simple
Rule
20 %
Small
Seen
52.89
54.30
74.54
60.93
61.67
76.23
86.27
80.29
Half
42.40
47.35
59.14
44.11
54.80
61.43
82.98
84.11
Unseen
31.66
45.84
51.54
38.67
47.91
54.75
82.00
84.88
Medium
Seen
68.32
56.34
82.15
78.04
78.78
82.68
86.27
80.29
Half
53.57
51.64
64.33
52.85
66.21
65.12
82.98
84.11
Table 3: F1 on the German version per corner-case ratio (CC), development set size, and test set. Bold: best, underline: second best per variant. Zero-shot GPT-5.2 is independent of the development set size.
Supervised
GPT-5.2 zero-shot
CC
Size
Test
WordCooc
Magellan
RoBERTa
R-SupCon
HierGAT
Ditto
Simple
Rule
20 %
Small
Seen
+4.23
−2.55
+2.56
+2.44
+12.41
−5.71
+0.20
−2.21
Half
+2.19
−1.58
+2.05
−0.08
+9.70
−3.40
−1.43
−0.48
Unseen
+5.72
−1.73
+4.43
−0.11
+11.10
−2.31
−0.84
+0.15
Medium
Seen
−2.05
−2.43
+1.38
−0.12
−1.54
+3.35
+0.20
−2.21
Half
−1.47
−3.34
+0.91
−0.92
−1.23
+4.53
−1.43
−0.48
Table 4: F1 difference between the English and the German version (EN − DE) per variant, positive values favour English. Adding a cell to Table 3 gives the English F1. Summary rows: mean by development set size, overall mean (bold) and median, and count of variants favouring English.
German
English
Size
Test
RoBERTa
XLM-R
RoBERTa
XLM-R
Small
Seen
47.91
57.54
56.50
57.51
Half
39.55
48.07
45.37
49.13
Unseen
37.31
43.31
42.07
45.23
Medium
Seen
64.39
68.66
69.59
71.77
Half
49.30
55.65
57.65
57.96
Table 5: F1 of RoBERTa and XLM-R on the 80 % corner-case variants of both language versions, mean over three seeds.
Records
Supervised
GPT-5.2
Category
Train
Test
WordCooc
Magellan
RoBERTa
R-SupCon
HierGAT
Ditto
Simple
Rule
Furniture & Living
11,946
2,760
36.88
38.27
49.46
44.85
48.35
50.66
71.25
77.08
Electronics & Computers
11,737
885
49.55
58.19
77.82
67.34
77.40
80.95
90.81
90.76
Clothing & Accessories
9,100
2,514
38.20
33.35
49.48
43.35
46.96
52.06
64.76
70.94
Tools & DIY
6,206
643
50.90
55.81
70.39
62.83
71.46
72.27
90.25
88.05
Toys & Baby
3,754
367
48.80
51.74
76.70
70.24
72.80
75.88
92.53
88.52
Table 6: Mean same-category F1 on German over nine test sets, seeds, and development sizes. Bold: best per category. Train/Test count record occurrences in pairs of the large 80 % corner-case training/Half-Seen test sets. Pet Supplies has no records in this test set and is omitted. Books, Films & Music has fewer than 50 test records.
DE-DE
Difference to DE-DE
Model
F1
DE-EN
EN-DE
EN-EN
Mixed
WordCooc
36.2
−6.0
−6.6
−10.5
−7.0
Magellan
33.8
−2.2
−3.0
−0.5
−1.2
RoBERTa
53.3
−7.2
−7.7
−5.3
−5.2
XLM-R
58.0
−4.8
−5.6
+0.3
−1.4
R-SupCon
43.1
−0.8
−1.8
−5.8
−2.2
Table 7: F1 on the five language variants of the 80 % corner-case Half-Seen test set and difference to DE-DE, computed from unrounded means. Supervised matchers are trained on German pairs only and report means over seeds, GPT-5.2 means over three repetitions.
Benchmark
Domain
Size
Multiling.
Sizes
Corner c.
Unseen
Cross-lang.
Abt-Buy [ KTR10 ]
electronics, other goods
9,575 pairs
–
–
–
–
–
Walmart-Amazon [ Ko16 ]
electronics
10,242 pairs
–
–
–
–
–
WDC Products [ PDB24 ]
electronics, other goods
11,715 offers
–
✓
✓
✓
–
Ember [ Wa22 ]
clothing and shoes
126,277 records
–
–
–
✓
–
Polish PM dataset [ Mo22 ]
drugstore, beverages
24,752 pairs
–
✓
–
–
–
ProMapCz [ MP23 ]
consumer products
1,495 pairs
–
–
–
–
–
Table 8: Product matching benchmarks. Multiling.: records in multiple languages. Sizes: multiple development set sizes. Corner c./Unseen: controlled corner-case/unseen-entity ratios. Cross-lang.: different languages within a pair.
Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model variant(reflecting different pretraining objectives), and model size, making it difficult to isolate the sources of performance gains. We address this issue through a controlled factorial study spanning three matcher architectures, three model variants and three model sizes from the Qwen3 family, and nine datasets, totaling 1,215 fine-tuning runs. We also evaluate cross-dataset transferability and computational cost. Our results show that model variant is critical for bi-encoders: embedding-oriented variants provide stronger initialization and more favorable representation geometry predictive of downstream matching performance. Cross-encoders retain a consistent advantage over bi-encoders because they jointly encode record pairs rather than representing each record independently, although larger models partially narrow this gap. Generative matchers do not universally outperform cross-encoders. Instead, their advantages concentrate under distribution shift, including subtle unseen differences in record schemas and cross-dataset transfer. We further find that larger models rely more heavily on shortcut learning and therefore do not necessarily perform better. These findings clarify the factors underlying performance differences across matcher architectures and motivate future research and benchmark designs that better disentangle architectural choices from model-level factors while explicitly evaluating distribution shift and cross-dataset transferability. We release our experimental results, code, training scripts, and evaluation data at https://github.com/Jantory/llm-trained-matcher.
Zeyu Zhang, Xue Li, Iacer Calixto +2
University of Amsterdam & AUMC · CWI · University of Amsterdam +1
Large language models (LLMs) achieve strong entity matching performance without task-specific training data, but applying them to large sets of candidate pairs is slow and costly. Matchers built on pretrained language models (PLMs), such as BERT, offer faster inference but require training data. We systematically study knowledge-distillation workflows in which an LLM teacher labels training pairs for a smaller student matcher. We vary pair selection, labeling budget, teacher model, correspondence post-processing, and student model across eight benchmarks, including unseen entities and non-English data. We compare students trained on machine-labeled data with matchers trained on the original benchmark training sets. In most cases, PLM-based matchers trained on LLM-labeled data perform similarly to those trained on benchmark sets. Pair selection matters most for small labeling budgets, where active learning is often most effective. An open-weight teacher trains competitive students, so distillation requires no closed-weight models. Compact PLM-based students compete with much larger LLM students on most tasks while requiring 34 to 459 times less inference time than direct LLM matching. On the two benchmarks with high shares of unseen products, PLM-based students substantially underperform their teachers, as do students trained on benchmark data. Under GPT-5.2 pricing, LLM labeling costs per training set average $5.86 to $8.11. These findings support knowledge distillation as a practical approach to reduce the effort of labeling task-specific training data while enabling efficient inference.
Aaron Steiner, Christian Bizer
Data and Web Science Group University of Mannheim Mannheim, Germany
We introduce a set of synthetic algorithmic tasks to detect cross-lingual gaps in the abilities of large language models. Our benchmark is commensurate across languages, since it requires models to perform the same underlying task in different languages; scalable, since each task can be generated at varying levels of complexity allowing it to be adapted to models with different capabilities; quantifiable, since every task admits an objective notion of correctness; and transparent, since tasks are generated from simple templates that can be readily audited for translation errors. Because our benchmark focuses on algorithmic tasks, differential performance is a sufficient -- but not necessary -- indicator of cross-lingual gaps. Nevertheless, we show through extensive experiments that our benchmark exposes persistent cross-lingual gaps in multiple state-of-the-art models.
Purvam Jain, Preethi Jyothi, Vihari Piratla +1
Google DeepMind · Indian Institute of Technology Bombay · International Centre for Theoretical Sciences, Tata Institute of Fundamental Research