Billiger.de Products: A Bilingual Entity Matching Benchmark
Authors: Aaron Steiner, Ksenia Elagin, Ralph Peeters, Johannes Knopp, Christian Bizer
Organizations: University of Mannheim, Data and Web Science Group, B6, 26, 68159 Mannheim, Germany · solute GmbH, Zeppelinstraße 15, 76185 Karlsruhe, Germany
Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces Billiger.de Products, a bilingual German and English entity matching benchmark covering thirteen consumer product categories, including difficult-to-handle categories such as clothing and furniture. The benchmark data originates from the German price comparison platform billiger.de. Following the design of WDC Products, the benchmark offers multiple variants that differ in the fraction of corner cases, the size of the development set, and the fraction of entities unseen during training. An aligned English translation of every offer keeps all pairs, splits, and labels fixed, while cross-language test sets combine German and English records within individual pairs. We validate the benchmark using six supervised matchers and zero-shot GPT-5.2 on both language versions and the cross-language test sets. The validation shows the difficulty of the benchmark. The comparison of the results on the English version of the benchmark to the results on the German version shows that most matchers score on average higher on the English version. The difference is largest for RoBERTa and HierGAT, while the zero-shot LLM runs are largely insensitive to the language. Comparing the F1 scores achieved by PLM-based matchers on the English version of Billiger.de Products with their performance on existing English-language benchmarks, such as WDC Products and Abt-Buy, shows that Billiger.de Products is more difficult than these benchmarks.
Figures & tables
Figure 2: Construction pipeline of the benchmark, from clustering to the aligned German and English releases.
Training
Validation
Test
CC
Size
Pos.
Neg.
Total
Pos.
Neg.
Total
Unseen
Pos.
Neg.
Total
20 %
Small
485
1,979
2,464
491
1,993
2,484
0 %
407
4,038
4,445
Medium
1,457
4,456
5,913
491
2,989
3,480
50 %
347
4,082
4,429
Large
13,100
14,012
27,112
491
3,984
4,475
100 %
360
4,089
4,449
50 %
Small
482
1,979
2,461
479
1,986
2,465
0 %
363
4,049
4,412
Medium
1,464
4,441
5,905
479
2,977
3,456
50 %
326
4,092
4,418
Table 1: Pair composition per corner-case ratio (CC), identical for both languages. Validation reports the Seen splits used for model selection. Training and validation vary by development set size, while test sets vary by unseen-product ratio.
Category
Train (%)
Test (%)
Category
Train (%)
Test (%)
Clothing & Accessories
20.85
28.56
Cosmetics & Drugstore
3.25
4.51
Furniture & Living
21.10
26.42
Health & Care
2.24
2.46
Electronics & Computers
20.53
12.51
Office Supplies
0.94
0.69
Tools & DIY
11.34
8.10
Food & Beverages
1.00
0.68
Sports & Leisure
6.23
6.22
Books, Films & Music
0.49
0.51
Toys & Baby
6.24
4.92
Pet Supplies
0.23
0.19
Table 2: Share of product offers per category in the training and test splits in percent, averaged over all variants.
Supervised
GPT-5.2 zero-shot
CC
Size
Test
WordCooc
Magellan
RoBERTa
R-SupCon
HierGAT
Ditto
Simple
Rule
20 %
Small
Seen
52.89
54.30
74.54
60.93
61.67
76.23
86.27
80.29
Half
42.40
47.35
59.14
44.11
54.80
61.43
82.98
84.11
Unseen
31.66
45.84
51.54
38.67
47.91
54.75
82.00
84.88
Medium
Seen
68.32
56.34
82.15
78.04
78.78
82.68
86.27
80.29
Half
53.57
51.64
64.33
52.85
66.21
65.12
82.98
84.11
Table 3: F1 on the German version per corner-case ratio (CC), development set size, and test set. Bold: best, underline: second best per variant. Zero-shot GPT-5.2 is independent of the development set size.
Supervised
GPT-5.2 zero-shot
CC
Size
Test
WordCooc
Magellan
RoBERTa
R-SupCon
HierGAT
Ditto
Simple
Rule
20 %
Small
Seen
+4.23
−2.55
+2.56
+2.44
+12.41
−5.71
+0.20
−2.21
Half
+2.19
−1.58
+2.05
−0.08
+9.70
−3.40
−1.43
−0.48
Unseen
+5.72
−1.73
+4.43
−0.11
+11.10
−2.31
−0.84
+0.15
Medium
Seen
−2.05
−2.43
+1.38
−0.12
−1.54
+3.35
+0.20
−2.21
Half
−1.47
−3.34
+0.91
−0.92
−1.23
+4.53
−1.43
−0.48
Table 4: F1 difference between the English and the German version (EN − DE) per variant, positive values favour English. Adding a cell to Table 3 gives the English F1. Summary rows: mean by development set size, overall mean (bold) and median, and count of variants favouring English.
German
English
Size
Test
RoBERTa
XLM-R
RoBERTa
XLM-R
Small
Seen
47.91
57.54
56.50
57.51
Half
39.55
48.07
45.37
49.13
Unseen
37.31
43.31
42.07
45.23
Medium
Seen
64.39
68.66
69.59
71.77
Half
49.30
55.65
57.65
57.96
Table 5: F1 of RoBERTa and XLM-R on the 80 % corner-case variants of both language versions, mean over three seeds.
Records
Supervised
GPT-5.2
Category
Train
Test
WordCooc
Magellan
RoBERTa
R-SupCon
HierGAT
Ditto
Simple
Rule
Furniture & Living
11,946
2,760
36.88
38.27
49.46
44.85
48.35
50.66
71.25
77.08
Electronics & Computers
11,737
885
49.55
58.19
77.82
67.34
77.40
80.95
90.81
90.76
Clothing & Accessories
9,100
2,514
38.20
33.35
49.48
43.35
46.96
52.06
64.76
70.94
Tools & DIY
6,206
643
50.90
55.81
70.39
62.83
71.46
72.27
90.25
88.05
Toys & Baby
3,754
367
48.80
51.74
76.70
70.24
72.80
75.88
92.53
88.52
Table 6: Mean same-category F1 on German over nine test sets, seeds, and development sizes. Bold: best per category. Train/Test count record occurrences in pairs of the large 80 % corner-case training/Half-Seen test sets. Pet Supplies has no records in this test set and is omitted. Books, Films & Music has fewer than 50 test records.
DE-DE
Difference to DE-DE
Model
F1
DE-EN
EN-DE
EN-EN
Mixed
WordCooc
36.2
−6.0
−6.6
−10.5
−7.0
Magellan
33.8
−2.2
−3.0
−0.5
−1.2
RoBERTa
53.3
−7.2
−7.7
−5.3
−5.2
XLM-R
58.0
−4.8
−5.6
+0.3
−1.4
R-SupCon
43.1
−0.8
−1.8
−5.8
−2.2
Table 7: F1 on the five language variants of the 80 % corner-case Half-Seen test set and difference to DE-DE, computed from unrounded means. Supervised matchers are trained on German pairs only and report means over seeds, GPT-5.2 means over three repetitions.
Benchmark
Domain
Size
Multiling.
Sizes
Corner c.
Unseen
Cross-lang.
Abt-Buy [ KTR10 ]
electronics, other goods
9,575 pairs
–
–
–
–
–
Walmart-Amazon [ Ko16 ]
electronics
10,242 pairs
–
–
–
–
–
WDC Products [ PDB24 ]
electronics, other goods
11,715 offers
–
✓
✓
✓
–
Ember [ Wa22 ]
clothing and shoes
126,277 records
–
–
–
✓
–
Polish PM dataset [ Mo22 ]
drugstore, beverages
24,752 pairs
–
✓
–
–
–
ProMapCz [ MP23 ]
consumer products
1,495 pairs
–
–
–
–
–
Table 8: Product matching benchmarks. Multiling.: records in multiple languages. Sizes: multiple development set sizes. Corner c./Unseen: controlled corner-case/unseen-entity ratios. Cross-lang.: different languages within a pair.