COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages
Organizations: IIT Patna · IIT Delhi · IIT Guwahati · IIIT Delhi · IGDTUW · MIT-MAHE
Abstract
Machine translation (MT) for Indian languages remains constrained by the limited availability of high-quality, Indic-centric parallel corpora and evaluation benchmarks. Existing multilingual resources are largely constructed from English-pivot content and often fail to capture the linguistic diversity, cultural complexity, and domain-specific characteristics of Indian languages. We present COILD, an Indic-centric parallel corpus comprising over 1.16 million human-translated and human-verified sentence pairs, covering 20 Indian language pairs across the Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic language families. The corpus is built entirely from original Indian language sources collected from licensed repositories spanning eight domains with direct real-world applicability. Furthermore, we introduce a domain-centric benchmark comprising 2,000 expert-verified sentences to enable consistent multilingual and cross-lingual evaluation across Indian language pairs. To validate the effectiveness of COILD, we fine-tune two representative multilingual neural machine translation models, IndicTrans2-Distilled and NLLB-200. Experimental results demonstrate consistent improvements across language pairs, domains, automatic evaluation metrics, and human evaluation, highlighting the effectiveness of high-quality Indic-centric supervision. COILD provides a valuable training and evaluation resource for advancing multilingual machine translation and future multilingual language models for Indian languages.
Figures & tables
| Source | Target | Sentence Pairs | Token Count | Vocabulary | ||
|---|---|---|---|---|---|---|
| Source | Target | Source | Target | |||
| Hindi | Assamese | 94,872 | 19,86,178 | 15,90,680 | 76,956 | 1,18,170 |
| Hindi | Bengali | 73,066 | 14,89,733 | 11,76,804 | 58,937 | 79,616 |
| Hindi | Bodo | 47,599 | 9,88,036 | 7,62,802 | 47,698 | 74,469 |
| Hindi | Dogri | 58,896 | 11,93,781 | 12,09,040 | 53,385 | 56,814 |
| Hindi | Gujarati | 96,612 | 20,22,924 | 16,84,920 | 78,395 | 1,15,035 |
| Hyperparameter | IndicTrans2 | NLLB-200 |
|---|---|---|
| Epochs | 5 | 5 |
| Learning Rate | ||
| LR Scheduler | Cosine | Linear |
| Warm-up | 1% | 3% |
| Optimizer | AdamW (8-bit) | AdamW (8-bit) |
| Effective Batch Size | 64 | 64 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| IndicTrans2 | NLLB | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Direction | Adequacy | Fluency | Adequacy | Fluency | ||||||||
| Baseline | Finetune | Baseline | Finetune | Baseline | Finetune | Baseline | Finetune | |||||
| Assamese Hindi | 2.16 | 2.49 | 0.33 | 3.01 | 3.35 | 0.34 | 2.52 | 2.71 | 0.19 | 3.42 | 3.54 | 0.12 |
| Hindi Assamese | 2.19 | 2.12 | -0.07 | 3.03 | 2.82 | -0.21 | 2.02 | 2.32 | 0.30 | 2.66 | 2.80 | 0.14 |
| Bengali Hindi | 3.86 | 4.22 | 0.36 | 4.15 | 4.44 | 0.29 | 3.43 | 3.89 | 0.47 | 3.52 | 4.02 | 0.49 |
| Hindi Bengali | 3.71 | 4.09 | 0.38 | 3.99 | 4.27 | 0.28 | 3.50 | 3.85 | 0.35 | 3.56 | 4.02 | 0.46 |