cs.CLSep 28, 2026

Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation

Authors: Thodsaporn Chay-intr, Krittapad Harnchang, Mahannop, Thabua, Kobkrit Viriyayudhakorn, Thanaruk Theeramunkong

Organizations: iApp Technology, Thailand · Intelligent Informatics and Service Innovation Research Center, Thailand · Artificial Intelligence Entrepreneur Association of Thailand (AIEAT), Thailand · Sirindhorn International Institute of Technology, Thammasat University, Thailand

Abstract

Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following

    May 27, 2026Mingrui Sun, Mao Zheng, Zheng Li +1Multilingual BenchmarkInstruction

  2. Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following

    Sep 23, 2026Niklas Scholz, David Thulke, Abdallah Nasir +3Large Language Model Fine-TuningMachine Translation

  3. Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

    Aug 11, 2026Chris Han, Pengzhi Gao, Pei Fu +1Post-Training