BaatCheet: A Multilingual Corpus for Dialogue Translation in Indian Languages
Organizations: Language Technologies Research Centre, IIIT Hyderabad, India
Abstract
Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks focus on sentence-level or formal text, making it difficult to assess translation quality of the dialogue phenomena. In this work, we introduce BaatCheet, a multilingual dialogue corpus named after the Hindi term for conversation or chitchat, containing approximately 49,000 dialogues for dialogue translation across five translation directions. We fine-tune five open-source LLMs across seven training data configurations and find that fine-tuning yields substantial gains over zero- and few-shot baselines. To comprehensively evaluate dialogue translation quality, we employ multiple evaluation strategies, including automatic metrics, LLM-as-judge, and human assessments using an SQM-guided Direct Assessment (DA) Protocol.
Figures & tables
| BaatCheet Corpus | #Dialogues | #Utterances | #Sentences | Avg Dlg Len | Avg Utt Len | Avg Sent Len | |||
| Src | Tgt | Src | Tgt | Src | Tgt | ||||
| BaatCheet_Syn_X2IL | 42,148 | 301,278 | 533,102 | 533,494 | 8.8 | 14.8 | 11.8 | 8.7 | 6.9 |
| BaatCheet_Syn_IL2X | 6,096 | 73,666 | 135,894 | 136,616 | 12.3 | 14.6 | 11.8 | 8.4 | 6.9 |
| BaatCheet_Human_Train | 510 | 8,405 | 16,455 | 16,077 | 17.3 | 17.6 | 14.1 | 9.1 | 7.5 |
| BaatCheet_Test | 658 | 8,904 | 16,813 | 16,518 | 13.7 | 16.9 | 14.1 | 8.9 | 7.5 |
| TOTAL | 49,412 | 392,253 | 702,264 | 702,705 | 13.0 | 16.0 | 13.0 | 8.8 | 7.2 |
| Baselines | Synthetic-Only | Human-Only | Dual-Hybrid | Multi-Hybrid | |||||
| Model | Zero-Shot | Few-Shot | Syn_X2IL | Syn_IL2X | HT | Syn_X2IL+HT | Syn_IL2X+HT | Syn_Both+HT | Syn_Both+3 HT |
| (a) COMET (Reference-Based) | |||||||||
| Gemma3-4B | 0.5024 | 0.6410 | 0.8076 | 0.7095 | 0.8324 | 0.8192 | 0.8223 | 0.7397 | 0.6087 |
| Llama-3.2-3B | 0.5317 | 0.5864 | 0.7802 | 0.7279 | 0.7139 | 0.7410 | 0.7076 | 0.7647 | 0.7271 |
| Llama-3.1-8B | 0.7113 | 0.6882 | 0.8023 | 0.7418 | 0.7871 | 0.7197 | 0.7210 | 0.7930 | 0.8461 |
| Qwen3-4B | 0.6797 | 0.6601 | 0.8373 | 0.7372 | 0.7959 | 0.8063 | 0.7370 | 0.7695 | 0.8288 |
| GPT Judge | Gemini Judge | Human Judge | |||||||||||
| Utt-Level | Dia-Level Factors | Utt-Level | Dia-Level Factors | Utt-Level | Dia-Level Factors | ||||||||
| Model | Configuration | Acc | Nat | Inf | Coh | Acc | Nat | Inf | Coh | Acc | Nat | Inf | Coh |
| Gemma | Syn_X2IL+HT | 48.58 | 42.79 | 48.82 | 46.66 | 47.94 | 45.95 | 48.86 | 47.69 | 49.98 | 46.45 | 49.63 | 47.46 |
| Human_Train | 50.31 | 48.99 | 50.52 | 50.92 | 50.54 | 51.62 | 52.25 | 52.79 | 51.52 | 50.20 | 50.15 | 50.67 | |
| Llama-3B | Syn_X2IL+HT | 43.14 | 39.55 | 46.76 | 37.47 | 41.31 | 37.49 | 41.81 | 37.13 | 43.22 | 43.11 | 49.23 | 42.31 |
| Syn_X2IL | 44.41 | 43.64 | 41.54 | 39.82 | 43.96 | 41.95 | 40.78 | 40.50 | 43.35 | 44.88 | 46.25 | 43.91 | |
| GPT vs H | Ge vs Hu | GPT vs Ge | ||||
| Factor | ||||||
| Nat | 0.552 | 0.370 | 0.625 | 0.428 | 0.738 | 0.540 |
| Inf | 0.343 | 0.232 | 0.332 | 0.209 | 0.391 | 0.266 |
| Coh | 0.561 | 0.364 | 0.647 | 0.442 | 0.743 | 0.532 |
| Utt | 0.702 | 0.498 | 0.697 | 0.525 | 0.781 | 0.597 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Data Source – Src. Lang | Description |
| Hinglish-Chat – Hin | A synthetically generated Hinglish conversational dataset Khatri (2024) featuring everyday-life dialogues. The dataset is modified by LLM prompting to “Remove unnatural noun and pronoun usage from this dialogue, but it still has to make sense and write it in Hindi (Devanagari) script.” |
| MuTual (Multi-Turn Dialogue Reasoning) – Eng | A Chinese high school English listening comprehension dataset Cui et al. (2020) , adapted, preprocessed, and manually annotated to create a multi-turn dialogue reasoning corpus aimed at evaluating dialogue models’ contextual understanding and reasoning capabilities. |
| DailyTalk – Eng | A conversational speech dataset for TTS systems developed by sampling and modifying dialogues from the open-domain DailyDialog corpus Li et al. (2017) . Dialogues from DailyTalk Lee et al. (2022) are used here, annotated with speaker identifiers to capture natural conversational flow. |
| News Articles – Eng, Hin, Tam, Tel | News articles Mujadia and Sharma (2025) from diverse domains are used to generate informal dialogues using LLM prompting: “Generate 10 informal dialogues between two friends on this topic. Keep them short, natural, conversational and direct, and conclude meaningfully.” |
| Short Stories – Tam, Tel | Short stories from online sources such as tamil_stories 12 12 12 https://huggingface.co/datasets/aitamilnadu/tamil_stories for Tamil and the Chandamama Kathalu Dataset 13 13 13 https://code.swecha.org/telugu-ai/chandamama-kathalu-dataset for Telugu are used. Each story is prompted as: “You are given a [language] story text. Generate one informal dialogue between two [language] friends discussing the story in detail in a natural, code-mixed (with English) tone in [language]. The dialogue should be conversational and informal like spoken language. Reach an opinion about the story.” |
| Datasource | Total Dialogues | Edited Dialogues | Dialogue Edit Rate (%)↓ | Total Utts | Edited Utts | Utt Edit Rate(%)↓ | TER↓ (%) | CER↓ (%) |
| News-Eng † | 50 | 4 | 8.00 | 516 | 4 | 0.78 | 0.31 | 0.25 |
| Hinglish † | 49 | 49 | 100.00 | 967 | 638 | 65.98 | 21.99 | 19.60 |
| News-Hin † | 49 | 34 | 69.39 | 492 | 139 | 28.25 | 3.96 | 2.77 |
| News-Tam | 48 | 27 | 56.25 | 344 | 61 | 17.73 | 4.81 | 2.94 |
| News-Tel | 48 | 38 | 79.17 | 361 | 167 | 46.26 | 17.87 | 11.91 |
| Stories-Tam | 50 | 48 | 96.00 | 777 | 330 | 42.47 | 9.68 | 6.21 |
| Direction | Data Source | #Dialogues | #Utterances | #Sentences (Src / Tgt) | Avg Dlg Len | Avg Utt Len (Src / Tgt) | Avg Sent Len (Src / Tgt) |
| (a) BaatCheet_Syn_X2IL ( Direction: English / Hindi Indic Language) | |||||||
| Eng Hin | DailyTalk | 2,446 | 22,152 | 32,582 / 32,312 | 9.06 | 11.28 / 11.41 | 7.81 / 7.94 |
| MuTual | 8,346 | 38,332 | 68,560 / 68,594 | 4.59 | 16.35 / 17.56 | 9.29 / 9.93 | |
| News | 826 | 7,834 | 12,690 / 12,663 | 9.48 | 11.86 / 13.39 | 7.34 / 8.30 | |
| Total | 11,618 | 68,318 | 113,832 / 113,569 | 5.88 | 14.97 / 15.97 | 8.84 / 9.39 | |
| Eng Tam | DailyTalk | 2,446 | 22,152 | 32,582 / 32,585 | 9.06 | 11.28 / 8.52 | 7.81 / 5.90 |
| Open-Source LLMs | Dedicated MT Systems | Proprietary LLM | |||||||
| Lang Pair | Llama-3B | Llama-8B | Qwen3-4B | Gemma3-4B | Sarvam-T | GMT | IndicTrans2 | BhashaVerse | GPT-4o-mini |
| (a) Reference-Based COMET | |||||||||
| Eng Hin | 0.7282 | 0.7546 | 0.6989 | 0.7881 | 0.8038 | 0.7957 | 0.7524 | 0.7543 | 0.7857 |
| Eng Tam | 0.5745 | 0.7141 | 0.4945 | 0.8109 | 0.7725 | 0.8623 | 0.8079 | 0.8100 | 0.8393 |
| Eng Tel | 0.5964 | 0.6257 | 0.5023 | 0.7644 | 0.8056 | 0.8391 | 0.8296 | 0.8210 | 0.8255 |
| Hin Tam | 0.5054 | 0.5675 | 0.5079 | 0.7998 | 0.6541 | 0.8791 | 0.8010 | 0.8554 | 0.8686 |
| Reference-Based COMET | Reference-Free COMET-QE | ||||||||
| Lang Pair | Data Source | GMT | GPT | IndicTrans2 | BhashaVerse | GMT | GPT | IndicTrans2 | BhashaVerse |
| (a) Human_Train_X2IL ( Direction) | |||||||||
| Eng Hin | DailyTalk | 0.7628 | 0.7506 | 0.7350 | 0.7257 | 0.8587 | 0.8474 | 0.8417 | 0.8297 |
| MuTual | 0.8059 | 0.8064 | 0.7526 | 0.7724 | 0.8597 | 0.8492 | 0.8329 | 0.8440 | |
| News-gen | 0.8185 | 0.8000 | 0.7696 | 0.7649 | 0.8655 | 0.8527 | 0.8542 | 0.8544 | |
| AVG | 0.7957 | 0.7857 | 0.7524 | 0.7543 | 0.8613 | 0.8498 | 0.8430 | 0.8427 | |
| (a) Utt-Avg Strategy (Dialogue Counts, ) | (b) Utt-High Strategy (Utterance Turn Counts, ) | |||||||
| X2IL Direction ( ) | IL2X Direction ( ) | X2IL Direction ( ) | IL2X Direction ( ) | |||||
| MT Engine | COMET-QE | COMET | COMET-QE | COMET | COMET-QE | COMET | COMET-QE | COMET |
| GMT | 320 | 320 | 199 | 79 | 3,506 | 3,449 | 3,210 | 2,811 |
| GPT-4o-mini | 154 | 127 | 294 | 431 | 2,690 | 2,541 | 3,720 | 4,725 |
| IndicTrans2 | 32 | 37 | 16 | 0 | 1,154 | 1,186 | 1,017 | 573 |
| BhashaVerse | 4 | 26 | 1 | 0 | 1,077 | 1,251 | 480 | 318 |
| (a) DOC-COMET Metric Validation | (b) COMTAIL Metric Validation | |||||||
| Selection via COMET | Selection via COMET-QE | Selection via COMET | Selection via COMET-QE | |||||
| Lang Pair | Utt-Avg | Utt-High | Utt-Avg | Utt-High | Utt-Avg | Utt-High | Utt-Avg | Utt-High |
| (a) Human_Train_X2IL( ) | ||||||||
| Eng Hin | 0.9049 | 0.9041 | 0.9057 | 0.9082 | 0.9177 | 0.9189 | 0.9213 | 0.9263 |
| Eng Tam | 0.9012 | 0.9018 | 0.9033 | 0.9067 | 0.8198 | 0.8209 | 0.8270 | 0.8466 |
| Eng Tel | 0.8932 | 0.8958 | 0.8958 | 0.9004 | 0.9277 | 0.9305 | 0.9320 | 0.9371 |
| Translation Direction | Utt-Avg | Utt-High |
| X2IL Direction ( ) | 123 (55.9%) | 90 (40.9%) |
| IL2X Direction ( ) | 128 (58.2%) | 81 (36.8%) |
| Total | 251 (57.0%) | 171 (38.9%) |
| Hyperparameter | Value |
| Number of epochs | 1 |
| Batch size | 2 |
| Gradient accumulation step | 4 |
| Learning rate | 1e-4 |
| Maximum sequence length | 4096 |
| LoRA rank | 32 |
| Model | LP | Zero-Shot | Few-Shot | Syn_X2IL | Syn_IL2X | HT | Syn_X2IL+ HT | Syn_IL2X+ HT | Syn_Both +HT | Syn_Both +3*HT |
| Llama-3B | Eng-Hin | 0.496 | 0.558 | 0.751 | 0.748 | 0.748 | 0.747 | 0.736 | 0.712 | 0.604 |
| Eng-Tam | 0.539 | 0.572 | 0.775 | 0.699 | 0.672 | 0.733 | 0.670 | 0.769 | 0.727 | |
| Eng-Tel | 0.523 | 0.556 | 0.784 | 0.694 | 0.699 | 0.711 | 0.694 | 0.786 | 0.759 | |
| Hin-Tam | 0.567 | 0.642 | 0.790 | 0.753 | 0.723 | 0.764 | 0.730 | 0.778 | 0.786 | |
| Hin-Tel | 0.533 | 0.605 | 0.801 | 0.746 | 0.728 | 0.751 | 0.708 | 0.779 | 0.760 | |
| AVGs | 0.532 | 0.586 | 0.780 | 0.728 | 0.714 | 0.741 | 0.708 | 0.765 | 0.727 |
| Model | LP | Zero-Shot | Few-Shot | Syn_X2IL | Syn_IL2X | HT | Syn_X2IL+ HT | Syn_IL2X+ HT | Syn_Both +HT | Synt _Both +3*HT |
| Llama-3B | Eng-Hin | 0.675 | 0.495 | 0.710 | 0.700 | 0.705 | 0.705 | 0.688 | 0.661 | 0.507 |
| Eng-Tam | 0.556 | 0.552 | 0.682 | 0.577 | 0.587 | 0.675 | 0.584 | 0.688 | 0.659 | |
| Eng-Tel | 0.590 | 0.525 | 0.713 | 0.581 | 0.601 | 0.650 | 0.608 | 0.710 | 0.699 | |
| Hin-Tam | 0.499 | 0.531 | 0.555 | 0.602 | 0.588 | 0.590 | 0.593 | 0.617 | 0.729 | |
| Hin-Tel | 0.523 | 0.474 | 0.572 | 0.597 | 0.563 | 0.589 | 0.578 | 0.635 | 0.700 | |
| AVGs | 0.568 | 0.515 | 0.647 | 0.611 | 0.609 | 0.642 | 0.610 | 0.662 | 0.658 |
| Model | LP | Zero-Shot | Few-Shot | Syn_X2IL | Syn_IL2X | HT | Syn_X2IL+ HT | Syn_IL2X+ HT | Synt _Both +HT | Syn_Both +3*HT |
| Llama-3B | Eng-Hin | 0.394 | 0.498 | 0.755 | 0.715 | 0.723 | 0.717 | 0.713 | 0.698 | 0.555 |
| Eng-Tam | 0.417 | 0.434 | 0.615 | 0.524 | 0.513 | 0.558 | 0.506 | 0.610 | 0.553 | |
| Eng-Tel | 0.301 | 0.315 | 0.700 | 0.524 | 0.515 | 0.536 | 0.497 | 0.701 | 0.658 | |
| Hin-Tam | 0.390 | 0.430 | 0.621 | 0.549 | 0.529 | 0.567 | 0.535 | 0.609 | 0.604 | |
| Hin-Tel | 0.306 | 0.408 | 0.716 | 0.606 | 0.581 | 0.621 | 0.553 | 0.681 | 0.648 | |
| AVGs | 0.362 | 0.417 | 0.681 | 0.584 | 0.572 | 0.600 | 0.561 | 0.660 | 0.604 |
| 7-Scale SQM Rubric | Eng–Hin | Eng–Tam | Eng–Tel | Hin–Tam | Hin–Tel |
| (a) Utterance-Level Accuracy | |||||
| 00–15: Completely unusable | 0 | 0 | 0 | 0 | 0 |
| 16–30: Very poor translation | 0 | 0 | 0 | 0 | 0 |
| 31–45: Meaning partially preserved | 0 | 0 | 0 | 0 | 0 |
| 46–60: Meaning correct but unnatural | 0 | 1 | 0 | 1 | 1 |
| 61–75: Mostly accurate and somewhat natural | 6 | 25 | 23 | 13 | 11 |
| GPT vs Human | Gemini vs Human | GPT vs Gemini | |||||
| Factor | Model | ||||||
| Naturalness | Gemma3-4B | 0.544 | 0.361 | 0.574 | 0.395 | 0.687 | 0.504 |
| Llama-3.2-3B | 0.540 | 0.382 | 0.606 | 0.427 | 0.758 | 0.548 | |
| Llama-3.1-8B | 0.594 | 0.422 | 0.665 | 0.486 | 0.788 | 0.605 | |
| Qwen3-4B | 0.543 | 0.321 | 0.639 | 0.397 | 0.734 | 0.514 | |
| Sarvam-T | 0.539 | 0.367 | 0.639 | 0.435 | 0.723 | 0.529 | |