Organizations: iApp Technology, Thailand · Intelligent Informatics and Service Innovation Research Center, Thailand · Artificial Intelligence Entrepreneur Association of Thailand (AIEAT), Thailand · Sirindhorn International Institute of Technology, Thammasat University, Thailand
Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.
Figures & tables
Figure 1: The two-phase RGDC pipeline. Phase 1 uses IFD scoring to cherry-select 1.75M instances from the 17.85M-record English-Thai parallel pool. Phase 2 extracts reference-grounded constraints and keeps only all-pass candidates, yielding the 1.97M-record Grounded dataset that trains the ChindaMT family.
Category
Example
Content
“include the term ‘clean energy”’
Numerical
“use exactly one sentence”
Stylistic
“use formal Thai without slang”
Format
“return only the translation, no quotes”
Linguistic
“avoid English contractions”
Table 1: Example constraints for the five categories used by the Phase 2A extractor.
Sub-step
Records
Phase 1: Cherry selection
1C. Cherry-selected (top-decile)
1,751,962
Phase 2: Constraint augmentation
2A. Constraint extractions
1,751,962
2B. Evaluation questions
1,747,959
2C. Augmented prompts (2 variants)
3,495,704
Table 2: Record counts at each RGDC sub-step, starting from a 17.85M parallel pool. Sub-steps 1B and 2D do not change record counts.
Model
Comparator
Plain (shared)
Plain (own)
Constrained
en → th
th → en
Mean
en → th
th → en
Mean
en → th
th → en
Mean
CoreEval
ChindaMT-4B
MiLMMT-46-4B
88.8
83.7
86.3
62.8
42.1
53.3
89.9
88.7
89.5
TranslateGemma-4B
88.7
75.1
81.0
80.1
54.5
66.7
86.4
91.0
87.2
Typhoon-Translate-4B
64.9
58.5
61.8
62.1
53.4
57.8
67.4
71.0
68.4
HY-MT-1.5-7B
68.2
83.0
74.0
68.2
59.3
62.9
96.6
99.5
97.9
GemmaX2-28-9B
94.7
92.8
93.6
66.2
45.0
56.6
94.4
93.5
94.1
Table 3: Main pairwise LC% on CoreEval and BroadEval. (shared) uses a uniform prompt scaffold, (own) lets each baseline use its own template. Bold marks Means above 50. Shared-prompt comparisons where the comparator returns almost no target-language output (GemmaX2 BroadEval, MM-46-4B BroadEval, and Constrained BroadEval for MM-46-1B at the 2B and 0.8B tiers) are omitted.
Model
CK ↑
DA ↑
MQM ↑
Quality Mean ↑
MtX ↓
MtX Mean ↓
en → th
th → en
en → th
th → en
en → th
th → en
en → th
th → en
FLORES-200
Larger [-1pt]4B–9B
MiLMMT-46-4B
83.88
84.98
87.94
91.90
97.01
98.23
90.7
2.14
1.97
2.06
TranslateGemma-4B
82.76
84.01
82.44
87.62
95.24
97.52
88.3
2.45
2.11
2.28
Typhoon-Translate-4B
83.62
83.97
87.52
87.79
97.23
97.74
89.6
2.14
2.26
2.20
HY-MT-1.5-7B
84.38
84.61
88.71
89.48
97.63
97.97
90.5
1.98
1.98
1.98
GemmaX2-28-9B
82.43
84.45
85.20
91.91
96.94
98.34
89.9
2.45
2.07
2.26
Table 4: External-metric translation quality on FLORES-200 (Wikipedia) and WMT24++ en-th (news), with CometKiwi (CK), GEMBA-DA (DA), and GEMBA-MQM (MQM) as 0 – 100 quality scores ( ↑ ); Mean ↑ is their unweighted six-value average. MetricX-24 (MtX, ↓ ) is the native error score on a 0 – 25 scale (lower better), kept separate with its own Mean ↓ over the two directions. Rows ordered by size within each tier, ChindaMT last. Bold marks the best per column within a tier, underline the second, ties bolded jointly.
Model
Comparator
Plain
Constr.
ChindaMT-4B
MM-46-4B
51.0
89.6
TG-4B
63.3
85.1
Typhoon-4B
56.1
65.5
HY-MT-7B
57.2
95.3
GemmaX2-9B
56.4
93.7
ChindaMT-2B
MM-46-1B
54.1
91.3
Table 5: Cross-judge ChindaMT WR% on CoreEval, averaged across the three open-weight judges. Plain uses each comparator’s own template, and values above 50 favor ChindaMT. Comparators are abbreviated, with full names in Table 11 . BroadEval, Plain (shared), and per-judge numbers in Appendix L .
Metric
Plain
Constrained
Overall
ChindaMT win rate
0.64
0.69
0.67
Pref. intensity
+0.39
+0.68
+0.53
Inter-rater κ
0.74
0.70
0.73
Human-vs-judge κ
0.18
0.62
0.41
Table 6: Human validation by three native Thai raters on 50 Plain and 50 Constrained items from the ChindaMT-4B against Typhoon-Translate-1.5-4B comparison, Plain with the baseline template and Constrained with the shared template. Win rate averages ChindaMT preference across raters (ties as half-wins), and intensity is the mean −2 to 2 rating. Inter-rater κ averages three pairs, and human-vs-judge κ compares against the human majority.
Components
Full recipe lift
Variant
Cherry
Filter
Rules
Plain
Constr.
Full
✓
✓
✓
—
—
no-cherry
×
✓
✓
+8.24
+0.69
no-all-pass
✓
×
✓
+1.09
+1.06
no-rules
✓
✓
×
+12.63
+9.53
Table 7: RGDC component ablations at 4B on CoreEval. Full recipe lift is the LC% margin over each ablation in pairwise judging, above the 50 tie. SE [2.1,2.5] .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Aspect
AutoIF
UltraIF
RGDC (ours)
Source data
36 hand-written seeds
ShareGPT prompts
Cherry-selected translation Q&A
Constraint origin
Self-instruct from seeds
Decomposed from prompts
Extracted from references
Verification
Python execution
LLM-as-judge
LLM-as-judge with explanations
Filter strictness
Score threshold
Rejection sampling
Every constraint must pass
Scale
10–25k
175k
1.97M
Aux. model
No
Yes (UltraComposer)
Yes (IFD scorer + aux LLM)
Appendix
Table 8: Comparison of RGDC with AutoIF and UltraIF across seven axes.
Corpus
Type
bible
Religious parallel ( Christodoulopoulos and Steedman, 2015 )
paracrawl
Web crawl ( Koehn, 2024 )
elrc
Government / news ( Tiedemann, 2012 )
hplt
Web crawl ( de Gibert et al., 2024 )
en-th-texts
Mixed ( kvush, 2024 )
opensubtitles
Film/TV subtitles ( Lison and Tiedemann, 2016 )
Appendix
Table 9: The 10 publicly available English-Thai parallel corpora (all en ↔ th) unified into the RGDC training pool, used by Phase 1 (cherry selection) and Phase 2 (constraint augmentation).
Slice
n
Δ CK
SE
hyp ≥ ref
Overall
20,000
+13.10
0.09
85.9%
en → th
10,000
+16.03
0.13
90.9%
th → en
10,000
+10.18
0.13
80.8%
Appendix
Table 10: CometKiwi audit of Phase 2C regeneration. Δ CK is the mean paired difference between the regenerated target and its source-corpus reference, and hyp ≥ ref is the share of records where the regenerated target scores at least as high.
Model
Size
ChindaMT (ours)
ChindaMT (Qwen3.5, 4B primary)
4B, 2B, 0.8B
ChindaMT-Qwen3-4B (cross-gen)
4B
External baselines
Typhoon-Translate-1.5
4B
Hunyuan-MT-1.5
7B, 1.8B
Appendix
Table 11: Models compared. TranslateGemma-4B and MiLMMT-46-4B are 4B-only comparators.
Sub-step
Temp.
Max tokens
2A. Constraint extraction
0.7
1024
2B. Evaluation questions
0
512
2C. Constrained responses
0.7
512
2D. Judge evaluation
0
1024
Appendix
Table 12: Temperature and max-token settings used by the auxiliary LLM at each Phase 2 sub-step. Identical across runs and bases.
Parameter
Value
Method
SFT (full parameters)
Epochs
1
Learning rate
2×10−5
Scheduler
inverse-square-root, 1% warmup
Weight decay
0.01
Optimizer
AdamW ( β1=0.9 , β2=0.999 )
Appendix
Table 13: Core ChindaMT-4B fine-tuning hyperparameters. The same recipe is used for all four variants.
Table 15: Per-direction pairwise win rate (WR%) by data source. CoreEval columns break the shared-prompt comparisons of Table 3 into directions (Plain and Constrained); external columns apply the same pairwise protocol (Section 4.6 ) to FLORES and WMT24++, which carry plain translation only. Dashes mark comparisons not run, as MiLMMT-46-1B has no external pairing. CoreEval columns report raw WR% per direction; external bootstrap SE is [1.3,1.6] .
Model
Comparator
Plain
Constr.
ChindaMT-4B
Qwen3.5-4B
55.7
62.2
ChindaMT-Qwen3-4B
Qwen3-4B
71.9
69.5
ChindaMT-2B
Qwen3.5-2B
77.1
74.2
ChindaMT-0.8B
Qwen3.5-0.8B
81.5
83.5
ChindaMT-4B
ChindaMT-2B
62.6
64.2
ChindaMT-2B
ChindaMT-0.8B
56.9
64.4
Appendix
Table 16: LC% on CoreEval. Upper block: each ChindaMT variant vs its own base; middle block: ChindaMT variants across parameter halving; last row: the cross-family Gemma-3-4B fine-tune (FT) vs its base. SE [1.8,2.4] , [2.2,2.4] , and 0.7/2.3 by block, the last small because the Plain cell is near ceiling. Under the primary, GPT-OSS-120B, and Llama-3.3-70B judges, the Gemma-3-4B row has Plain/Constrained WR% of 97.9/67.4 , 96.2/62.4 , and 97.9/67.0 ( n=400 ).
Model
Comparator
Plain (shared)
Plain (own)
Constrained
Pri
OSS
Llama
Pri
OSS
Llama
Pri
OSS
Llama
CoreEval
ChindaMT-4B
MiLMMT-46-4B
86.6
86.5
86.4
55.1
47.1
50.9
90.2
89.5
89.0
TranslateGemma-4B
79.6
77.8
78.6
66.5
59.9
63.5
85.6
83.2
86.4
Typhoon-Translate-4B
62.1
55.0
54.1
59.8
52.2
56.2
67.9
66.1
62.6
HY-MT-1.5-7B
72.0
61.6
65.5
62.1
53.0
56.6
96.5
94.8
94.8
GemmaX2-28-9B
93.8
93.0
92.2
59.4
54.1
55.6
95.1
92.8
93.1
Appendix
Table 17: Per-judge WR% for the three judges behind the Table 5 consensus. Direction-averaged Mean; values above 50 favor ChindaMT. n=400 , bootstrap SE [0.9,2.6] .
Dataset
Variant
CK
DA
MQM
Mean
FLORES-200
Full recipe
84.18
89.83
97.96
90.65
No-cherry
84.28
90.27
98.09
90.88
No-all-pass
84.19
89.68
97.92
90.60
No-rules
83.36
88.12
97.70
89.73
WMT24++
Full recipe
80.72
85.73
97.00
87.81
No-cherry
80.89
86.57
96.89
88.12
Appendix
Table 18: External-metric ablations at 4B, direction-averaged; ablations defined in Section 5.5 . Bold marks the best Mean per dataset.
Modern translation workflows demand more than semantic equivalence. Users routinely require models to preserve JSON or HTML schemas, honor curated glossaries, disambiguate with provided context, and match prescribed registers, often several at once. Conventional metrics such as BLEU and xCOMET capture semantic fidelity but provide little signal on constraint adherence, while general instruction following benchmarks ignore the cross-lingual nature of translation. We introduce \bench, a benchmark for multilingual translation instruction following covering seven languages, with 4,506 single-constraint and 2,838 multi-constraint items spanning six constraint dimensions and five compositional patterns with instructions issued in all seven languages. Constraints are split into a gating subset verified by deterministic checkers and a continuous subset scored by a rubric-based LLM judge, combined under a multiplicative rule that resists reward hacking. Evaluating 15 models reveals systematic gaps that prior protocols miss: Instruction following scales with size more sharply than translation quality, glossary and structured-format constraints dominate the difficulty gradient, and general instruction following rankings correlate only weakly with translation behavior. Our benchmark are available at https://github.com/Tencent-Hunyuan/Hy-MT2/tree/main/IFMTBench.
Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings transfer to machine translation (MT) fine-tuning and to MT-specific instruction following (MT-IF): instructions that modify a translation, such as formality, grammatical gender, and length control. We compare methods anchored to auxiliary data, to model outputs, and to the base model parameters, first in a screening study with Llama 3.2 1B Instruct, then on Llama 3.1 8B Instruct fine-tuned on bidirectional Arabic-English or Spanish-English data. Elastic Weight Consolidation preserves general capabilities best in both stages; on the 8B Spanish model the average score on general benchmarks drops 1.7 points versus 11.0 for standard fine-tuning, yet its scores for formality and grammatical gender control remain close to standard fine-tuning. Only data mixing with control-task examples preserves these controls, but its gains do not transfer to unseen prompts for the same task.
Niklas Scholz, David Thulke, Abdallah Nasir +3
AppTek GmbH, Aachen, Germany · Machine Learning and Human Language Technology, RWTH Aachen University, Germany · Applied Science Private University, Amman, Jordan
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification. We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0. Across 46 languages, the resulting models consistently improve translation quality over their SFT counterparts, outperform strong recent open baselines, including Seed-X, HY-MT2, and TranslateGemma, and achieve leading reference-free scores against evaluated proprietary systems such as Google Translate, Gemini 3 Pro, and GPT-5. We further investigate on-policy distillation (OPD) and find that it reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation. We release the models and code to facilitate future research.