Organizations: iApp Technology, Thailand · Intelligent Informatics and Service Innovation Research Center, Thailand · Artificial Intelligence Entrepreneur Association of Thailand (AIEAT), Thailand · Sirindhorn International Institute of Technology, Thammasat University, Thailand
Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.
Figures & tables
Figure 1: The two-phase RGDC pipeline. Phase 1 uses IFD scoring to cherry-select 1.75M instances from the 17.85M-record English-Thai parallel pool. Phase 2 extracts reference-grounded constraints and keeps only all-pass candidates, yielding the 1.97M-record Grounded dataset that trains the ChindaMT family.
Category
Example
Content
“include the term ‘clean energy”’
Numerical
“use exactly one sentence”
Stylistic
“use formal Thai without slang”
Format
“return only the translation, no quotes”
Linguistic
“avoid English contractions”
Table 1: Example constraints for the five categories used by the Phase 2A extractor.
Sub-step
Records
Phase 1: Cherry selection
1C. Cherry-selected (top-decile)
1,751,962
Phase 2: Constraint augmentation
2A. Constraint extractions
1,751,962
2B. Evaluation questions
1,747,959
2C. Augmented prompts (2 variants)
3,495,704
Table 2: Record counts at each RGDC sub-step, starting from a 17.85M parallel pool. Sub-steps 1B and 2D do not change record counts.
Model
Comparator
Plain (shared)
Plain (own)
Constrained
en → th
th → en
Mean
en → th
th → en
Mean
en → th
th → en
Mean
CoreEval
ChindaMT-4B
MiLMMT-46-4B
88.8
83.7
86.3
62.8
42.1
53.3
89.9
88.7
89.5
TranslateGemma-4B
88.7
75.1
81.0
80.1
54.5
66.7
86.4
91.0
87.2
Typhoon-Translate-4B
64.9
58.5
61.8
62.1
53.4
57.8
67.4
71.0
68.4
HY-MT-1.5-7B
68.2
83.0
74.0
68.2
59.3
62.9
96.6
99.5
97.9
GemmaX2-28-9B
94.7
92.8
93.6
66.2
45.0
56.6
94.4
93.5
94.1
Table 3: Main pairwise LC% on CoreEval and BroadEval. (shared) uses a uniform prompt scaffold, (own) lets each baseline use its own template. Bold marks Means above 50. Shared-prompt comparisons where the comparator returns almost no target-language output (GemmaX2 BroadEval, MM-46-4B BroadEval, and Constrained BroadEval for MM-46-1B at the 2B and 0.8B tiers) are omitted.
Model
CK ↑
DA ↑
MQM ↑
Quality Mean ↑
MtX ↓
MtX Mean ↓
en → th
th → en
en → th
th → en
en → th
th → en
en → th
th → en
FLORES-200
Larger [-1pt]4B–9B
MiLMMT-46-4B
83.88
84.98
87.94
91.90
97.01
98.23
90.7
2.14
1.97
2.06
TranslateGemma-4B
82.76
84.01
82.44
87.62
95.24
97.52
88.3
2.45
2.11
2.28
Typhoon-Translate-4B
83.62
83.97
87.52
87.79
97.23
97.74
89.6
2.14
2.26
2.20
HY-MT-1.5-7B
84.38
84.61
88.71
89.48
97.63
97.97
90.5
1.98
1.98
1.98
GemmaX2-28-9B
82.43
84.45
85.20
91.91
96.94
98.34
89.9
2.45
2.07
2.26
Table 4: External-metric translation quality on FLORES-200 (Wikipedia) and WMT24++ en-th (news), with CometKiwi (CK), GEMBA-DA (DA), and GEMBA-MQM (MQM) as 0 – 100 quality scores ( ↑ ); Mean ↑ is their unweighted six-value average. MetricX-24 (MtX, ↓ ) is the native error score on a 0 – 25 scale (lower better), kept separate with its own Mean ↓ over the two directions. Rows ordered by size within each tier, ChindaMT last. Bold marks the best per column within a tier, underline the second, ties bolded jointly.
Model
Comparator
Plain
Constr.
ChindaMT-4B
MM-46-4B
51.0
89.6
TG-4B
63.3
85.1
Typhoon-4B
56.1
65.5
HY-MT-7B
57.2
95.3
GemmaX2-9B
56.4
93.7
ChindaMT-2B
MM-46-1B
54.1
91.3
Table 5: Cross-judge ChindaMT WR% on CoreEval, averaged across the three open-weight judges. Plain uses each comparator’s own template, and values above 50 favor ChindaMT. Comparators are abbreviated, with full names in Table 11 . BroadEval, Plain (shared), and per-judge numbers in Appendix L .
Metric
Plain
Constrained
Overall
ChindaMT win rate
0.64
0.69
0.67
Pref. intensity
+0.39
+0.68
+0.53
Inter-rater κ
0.74
0.70
0.73
Human-vs-judge κ
0.18
0.62
0.41
Table 6: Human validation by three native Thai raters on 50 Plain and 50 Constrained items from the ChindaMT-4B against Typhoon-Translate-1.5-4B comparison, Plain with the baseline template and Constrained with the shared template. Win rate averages ChindaMT preference across raters (ties as half-wins), and intensity is the mean −2 to 2 rating. Inter-rater κ averages three pairs, and human-vs-judge κ compares against the human majority.
Components
Full recipe lift
Variant
Cherry
Filter
Rules
Plain
Constr.
Full
✓
✓
✓
—
—
no-cherry
×
✓
✓
+8.24
+0.69
no-all-pass
✓
×
✓
+1.09
+1.06
no-rules
✓
✓
×
+12.63
+9.53
Table 7: RGDC component ablations at 4B on CoreEval. Full recipe lift is the LC% margin over each ablation in pairwise judging, above the 50 tie. SE [2.1,2.5] .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Aspect
AutoIF
UltraIF
RGDC (ours)
Source data
36 hand-written seeds
ShareGPT prompts
Cherry-selected translation Q&A
Constraint origin
Self-instruct from seeds
Decomposed from prompts
Extracted from references
Verification
Python execution
LLM-as-judge
LLM-as-judge with explanations
Filter strictness
Score threshold
Rejection sampling
Every constraint must pass
Scale
10–25k
175k
1.97M
Aux. model
No
Yes (UltraComposer)
Yes (IFD scorer + aux LLM)
Appendix
Table 8: Comparison of RGDC with AutoIF and UltraIF across seven axes.
Corpus
Type
bible
Religious parallel ( Christodoulopoulos and Steedman, 2015 )
paracrawl
Web crawl ( Koehn, 2024 )
elrc
Government / news ( Tiedemann, 2012 )
hplt
Web crawl ( de Gibert et al., 2024 )
en-th-texts
Mixed ( kvush, 2024 )
opensubtitles
Film/TV subtitles ( Lison and Tiedemann, 2016 )
Appendix
Table 9: The 10 publicly available English-Thai parallel corpora (all en ↔ th) unified into the RGDC training pool, used by Phase 1 (cherry selection) and Phase 2 (constraint augmentation).
Slice
n
Δ CK
SE
hyp ≥ ref
Overall
20,000
+13.10
0.09
85.9%
en → th
10,000
+16.03
0.13
90.9%
th → en
10,000
+10.18
0.13
80.8%
Appendix
Table 10: CometKiwi audit of Phase 2C regeneration. Δ CK is the mean paired difference between the regenerated target and its source-corpus reference, and hyp ≥ ref is the share of records where the regenerated target scores at least as high.
Model
Size
ChindaMT (ours)
ChindaMT (Qwen3.5, 4B primary)
4B, 2B, 0.8B
ChindaMT-Qwen3-4B (cross-gen)
4B
External baselines
Typhoon-Translate-1.5
4B
Hunyuan-MT-1.5
7B, 1.8B
Appendix
Table 11: Models compared. TranslateGemma-4B and MiLMMT-46-4B are 4B-only comparators.
Sub-step
Temp.
Max tokens
2A. Constraint extraction
0.7
1024
2B. Evaluation questions
0
512
2C. Constrained responses
0.7
512
2D. Judge evaluation
0
1024
Appendix
Table 12: Temperature and max-token settings used by the auxiliary LLM at each Phase 2 sub-step. Identical across runs and bases.
Parameter
Value
Method
SFT (full parameters)
Epochs
1
Learning rate
2×10−5
Scheduler
inverse-square-root, 1% warmup
Weight decay
0.01
Optimizer
AdamW ( β1=0.9 , β2=0.999 )
Appendix
Table 13: Core ChindaMT-4B fine-tuning hyperparameters. The same recipe is used for all four variants.
Table 15: Per-direction pairwise win rate (WR%) by data source. CoreEval columns break the shared-prompt comparisons of Table 3 into directions (Plain and Constrained); external columns apply the same pairwise protocol (Section 4.6 ) to FLORES and WMT24++, which carry plain translation only. Dashes mark comparisons not run, as MiLMMT-46-1B has no external pairing. CoreEval columns report raw WR% per direction; external bootstrap SE is [1.3,1.6] .
Model
Comparator
Plain
Constr.
ChindaMT-4B
Qwen3.5-4B
55.7
62.2
ChindaMT-Qwen3-4B
Qwen3-4B
71.9
69.5
ChindaMT-2B
Qwen3.5-2B
77.1
74.2
ChindaMT-0.8B
Qwen3.5-0.8B
81.5
83.5
ChindaMT-4B
ChindaMT-2B
62.6
64.2
ChindaMT-2B
ChindaMT-0.8B
56.9
64.4
Appendix
Table 16: LC% on CoreEval. Upper block: each ChindaMT variant vs its own base; middle block: ChindaMT variants across parameter halving; last row: the cross-family Gemma-3-4B fine-tune (FT) vs its base. SE [1.8,2.4] , [2.2,2.4] , and 0.7/2.3 by block, the last small because the Plain cell is near ceiling. Under the primary, GPT-OSS-120B, and Llama-3.3-70B judges, the Gemma-3-4B row has Plain/Constrained WR% of 97.9/67.4 , 96.2/62.4 , and 97.9/67.0 ( n=400 ).
Model
Comparator
Plain (shared)
Plain (own)
Constrained
Pri
OSS
Llama
Pri
OSS
Llama
Pri
OSS
Llama
CoreEval
ChindaMT-4B
MiLMMT-46-4B
86.6
86.5
86.4
55.1
47.1
50.9
90.2
89.5
89.0
TranslateGemma-4B
79.6
77.8
78.6
66.5
59.9
63.5
85.6
83.2
86.4
Typhoon-Translate-4B
62.1
55.0
54.1
59.8
52.2
56.2
67.9
66.1
62.6
HY-MT-1.5-7B
72.0
61.6
65.5
62.1
53.0
56.6
96.5
94.8
94.8
GemmaX2-28-9B
93.8
93.0
92.2
59.4
54.1
55.6
95.1
92.8
93.1
Appendix
Table 17: Per-judge WR% for the three judges behind the Table 5 consensus. Direction-averaged Mean; values above 50 favor ChindaMT. n=400 , bootstrap SE [0.9,2.6] .
Dataset
Variant
CK
DA
MQM
Mean
FLORES-200
Full recipe
84.18
89.83
97.96
90.65
No-cherry
84.28
90.27
98.09
90.88
No-all-pass
84.19
89.68
97.92
90.60
No-rules
83.36
88.12
97.70
89.73
WMT24++
Full recipe
80.72
85.73
97.00
87.81
No-cherry
80.89
86.57
96.89
88.12
Appendix
Table 18: External-metric ablations at 4B, direction-averaged; ablations defined in Section 5.5 . Bold marks the best Mean per dataset.