Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.
Figures & tables
E5-large
LaBSE
Shared-offset proxy
75.4%
32.6%
Random-pair cosine
0.754±0.022
0.321
d (EN, ES)
0.081
0.145
d (EN, JA)
0.113
0.165
ES / paraphrase
7.4 ×
13.7 ×
JA / paraphrase
10.3 ×
15.6 ×
Table 1: Encoder diagnostics on 20,000 parallel products. The table uses the original full-English source for comparability across encoders. Probe chance is 33.3%.
Mix
Query translation
d=1
d=2
d=3
A
EN rewording
.901
.619
.396
A
ES
.676
.239
.081
A
JA
.643
.171
.037
B
EN rewording
.890
.605
.379
B
ES
.283
.055
.016
B
JA
.077
.002
.000
Table 2: E5 SID-prefix consistency on 4k held-out products, averaged over three seeds. A query succeeds at depth d when its first d codes match the SID assigned to the English listing.
Target distance
Translation
Isotropic
Product- directed
Spanish
.981
.756
.976
Japanese
.995
.887
.998
Table 3: Full SID mismatch under the E5 80/10/10 quantizer, using EN-Trunc as the source. Both controls match each translation’s cosine distance. The product-directed control uses a direction defined by another observed product embedding.
Encoder
Mix
Keff (L1)
Prefix@1
EN
ES
JA
ES
JA
E5
B
76
30
16
.283
.077
E5
C
49
42
42
.066
.048
LaBSE
B
101
91
90
–
–
Table 4: Effective first-level capacity on equal-size language pools and E5 first-code consistency with the English-indexed SID. Mixture B is 80/10/10; mixture C uses equal thirds.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
Change
L1 alive
Keff
Collision
V0
baseline
1
1
21%
V1
add centering
22
20
1.9%
V2
RVQ library, cosine
256
217
50%
V5
final configuration
256
102
4.3%
V7
reset every 2 epochs
8
6
99%
Appendix
Table 5: Representative collapse ablations on 20,000 items. V5 is used in the main experiments.
Spanish
Japanese
Mix
Rendering
Trans.
Iso.
Δ
Trans.
Iso.
Δ
B
surface
42.8
24.7
+18.1
77.7
47.2
+30.5
C
surface
46.9
24.7
+22.2
78.6
47.6
+31.0
B
NLLB
99.0
82.0
+17.0
99.8
91.4
+8.4
C
NLLB
99.5
86.1
+13.4
99.9
94.0
+5.9
Appendix
Table 6: full SID mismatch against the original isotropic control. These numbers are retained as measurements, not as evidence of language-selective quantizer amplification.
Metric
Qwen
NLLB
EN–ES distance
.0807
.0760
Distance / paraphrase
9.9 ×
9.3 ×
Jaccard@10
.453
.487
Keff (L1), BQ : EN/ES
50/11
49/17
Prefix@1 ES, AQ/BQ/CQ
.39/.24/.04
.34/.13/.003
Prefix@1 EN rewording
.78–.83
.76–.83
Appendix
Table 7: Full-corpus English–Spanish translator replication. The two-language mixtures are defined above and are separate from the main NLLB experiment.
Direction
Translation
Isotropic
Margin
EN → ES
.994
.794
+.200
ES → EN
.989
.717
+.272
EN → JA
.996
.895
+.100
JA → EN
.993
.902
+.090
Appendix
Table 8: Direction reversal under the E5 80/10/10 quantizer. The product-directed interpretation from the main paper also applies here.