Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation
Organizations: Indian Institute of Technology Palakkad · Mehta Family School of Data Science and Artificial Intelligence Indian Institute of Technology Palakkad, Kerala, India
Abstract
COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target into Latin script, which changes orthographic form while holding content and human ratings fixed. Script identity then accounts for 22.9% of native-script COMET variance, and agreement with annotators falls in all five languages studied. We trace the effect to the tokeniser and measure it with three label-free diagnostics. The bias is two faults, not one. Scores from different scripts occupy incompatible ranges, and within a single script the metric orders translations less accurately. No order-preserving transform of the score can repair the second fault. The first is removed exactly by COMET-QN, which maps the score distribution of each (language, script) pair onto a shared reference. Pooled agreement with annotators rises from 0.300 to 0.399, which is what makes scores from different scripts safe to place on one axis, and every within-language ordering is provably preserved. A regressor over parity features recovers a further 17.1% of the lost sensitivity. The remainder belongs to the encoder, and no post-processing can reach it. We therefore recommend publishing the normalised score, the three diagnostics, and the identity of the tokeniser they were computed against, so that a reader can tell how much of a score reflects translation quality and how much reflects the writing system.
Figures & tables
| (A) One translation, two scripts | ||
|---|---|---|
| “It was the final match for the All Blacks, who had already won the trophy two weeks ago.” | ||
| Condition | Human | COMET |
| MAL native script | ||
| Same text, romanised | ||
| (B) One source, two error-free targets | ||
| “Hydrogen ions are protons that had their electrons stripped off them.” | ||
| Language | COMET | TP | IP | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| nat | rom | nat | rom | nat | rom | |||||
| GUJ a | ||||||||||
| TAM | ||||||||||
| MAL | ||||||||||
| MAR b | ||||||||||
| HIN | ||||||||||
| Severity | Mean COMET | |
|---|---|---|
| Very Low | 69 | |
| Low | 86 | |
| Medium | 184 | |
| High | 258 | |
| Very High | 661 | |
| Segment-level severity– COMET: , | ||
| Corrector (transform class) | |||
|---|---|---|---|
| Raw romanised (identity) | |||
| Per-lang mean shift (location) | |||
| Per-lang affine / (loc.+scale) | |||
| Per-lang quantile map (monotone) | |||
| Per-lang isotonic (monotone) | |||
| Native score (ceiling) |
| Held-out | Recovered | |||
|---|---|---|---|---|
| GUJ | ||||
| TAM | ||||
| MAL | ||||
| MAR | ||||
| HIN | ||||
| Mean | ||||
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Lang. | N | Def | VL | L | M | H | VH |
|---|---|---|---|---|---|---|---|
| GUJ | |||||||
| TAM | |||||||
| MAL | |||||||
| MAR | |||||||
| HIN |
| Language | COMET | BLEURT | BERTScore | BLEU | ChrF | TER |
|---|---|---|---|---|---|---|
| GUJ | ||||||
| GUJ | ||||||
| TAM | ||||||
| TAM | ||||||
| MAL | ||||||
| MAL |
| Quantity | Native | Romanised |
|---|---|---|
| Segment-level proportions | ||
| GUJ | ||
| TAM | ||
| Metric | Ideal | Role |
|---|---|---|
| Fragmentation cost; over-fragmentation. | ||
| Semantic density per token; collapse. | ||
| Burden per unit information. Flag . | ||
| Distance from parity; Parity/Burden/Paradox zones. | ||
| Romanisation overhead; LP, EP defined in § 4.2 . |
| Lang. | Script | Family | mattr | DV/1k | IP–COMET | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| SPA | Latin | Romance | n/a | n/a | ||||||
| DEU | Latin | Germanic | n/a | n/a | ||||||
| GUJ | Gujarati | IA | ||||||||
| TAM | Tamil | Dravidian | ||||||||
| MAL | Malayalam | Dravidian | ||||||||
| MAR | Devanagari | IA |
| Lang. | LP | EP | Tax | EP% |
|---|---|---|---|---|
| GUJ | ||||
| TAM | ||||
| MAL | ||||
| MAR | ||||
| HIN |
| Lang. | Native | Romanised |
|---|---|---|
| GUJ | ||
| TAM | ||
| MAL | ||
| MAR | ||
| HIN |
| Setting | SBI | MATTR | Byte pr. | 1-tok% | IP–COMET | |
|---|---|---|---|---|---|---|
| ENG-SPA (SPA) | ||||||
| ENG-DEU (DEU) | ||||||
| Indic nat. | – | – | – | – | – | varies |
| Lang. | Script | MATTR nat | MATTR rom | Byte Prem. |
|---|---|---|---|---|
| GUJ | Gujarati | |||
| TAM | Tamil | |||
| MAL | Malayalam | |||
| MAR | Devanagari | |||
| HIN | Devanagari | |||
| SPA | Latin (Romance) | n/a |
| Raw | Rank | Position | Mapped |
|---|---|---|---|