Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection
Organizations: EPFL
Abstract
Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance. While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data. This work explores extending quality filtering to over 100 languages by proposing a multilingual adaptation approach that converts an existing English quality classifier into a multilingual variant. Our approach proposes training a small multi-layer perceptron on top of Transformer encoder-only model embeddings, using multilingual text as input and scores obtained from English classifiers applied to machine-translated text as labels. Our 1B, 3B and 8B scale experiments show that our approach maintains the downstream LLM benchmark performance of existing multilingual model-based filtering baselines, without harming regional and cultural knowledge benchmarks. To further evaluate cross-lingual generalization, we compare classifier scores of high-quality synthetic data and web samples, and the correlation of classifier scores with LLM-based ones, revealing that the classifier can learn the scoring criteria of its original English variant, even for languages not included in its training data.
Figures & tables
| GMMLU c | INCLUDE c | ARC | mARC | HS | mHS | XNLI | OBQA | XWG | Avg. | ||
| Classifier | Emb. Model | ||||||||||
| 1B Model, 100B Tokens from Scratch, Training and Evaluation on English + Top 20 Languages | |||||||||||
| mFW-edu | XLM-R | 0.2814 | 0.2944 | 0.4812 | 0.2936 | 0.5412 | 0.4034 | 0.4161 | 0.3800 | 0.6046 | 0.4106 |
| mmBERT | 0.2989 | 0.2913 | 0.4897 | 0.2937 | 0.5469 | 0.4026 | 0.4121 | 0.3740 | 0.6180 | 0.4141 | |
| mDCLM | XLM-R | 0.2845 | 0.2965 | 0.4802 | 0.2815 | 0.5651 | 0.4080 | 0.4227 | 0.3640 | 0.6277 | 0.4145 |
| mmBERT | 0.2848 | 0.2938 | 0.4554 | 0.2744 | 0.5454 | 0.4012 | 0.4142 | 0.3780 | 0.6314 | 0.4087 | |
| GMMLU c | INCLUDE c | ARC | mARC | HS | mHS | XNLI | OBQA | XWG | Avg. | |
| Model | ||||||||||
| 1B Model, 100B Tokens from Scratch, Training and Evaluation on English + Top 20 Languages | ||||||||||
| FW+FW2 | 0.2743 | 0.2817 | 0.3880 | 0.2586 | 0.5434 | 0.3887 | 0.4086 | 0.3180 | 0.6051 | 0.3852 |
| FWHQ+FW2HQ | 0.2893 | 0.2870 | 0.4461 | 0.2747 | 0.5641 | 0.4035 | 0.4214 | 0.3680 | 0.6308 | 0.4094 |
| FWHQ+FW2HQ + | 0.2952 | 0.2959 | 0.4505 | 0.2816 | 0.5762 | 0.4120 | 0.4116 | 0.3640 | 0.6221 | 0.4121 |
| mFW-edu (Ours) | 0.2989 | 0.2913 | 0.4897 | 0.2937 | 0.5469 | 0.4026 | 0.4121 | 0.3740 | 0.6180 | 0.4141 |
| Chinese | Japanese | French | Portuguese | Avg. | |
| Model | |||||
| 1B Model, 30B Tokens From Scratch, Monolingual Training and Evaluation | |||||
| FWHQ+FW2HQ + | 0.3578 | 0.3453 | 0.3938 | 0.3140 | 0.3527 |
| mFW-edu (Ours) | 0.3358 | 0.3653 | 0.4153 | 0.3230 | 0.3598 |
| mFW-edu (Ours) | 0.3541 | 0.3453 | 0.3771 | 0.3339 | 0.3526 |
| China | Japan | France | Brazil | Avg. (All Subsets) | |
| Model | |||||
| 8B Model, 63B Token Cooldown from 9.5T Token Checkpoint, Training and | |||||
| Evaluation on All Available Languages | |||||
| Pre-cooldown | 0.4915 | 0.6792 | 0.6429 | 0.4000 | 0.5764 |
| FWHQ+FW2HQ + | 0.5424 | 0.7925 | 0.6429 | 0.3600 | 0.6143 |
| mFW-edu (Ours) | 0.5763 | 0.7170 | 0.5714 | 0.4000 | 0.6164 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| GMMLU c | INCLUDE c | ARC | mARC | HS | mHS | XNLI | OBQA | XWG | Avg. | |
| Model | ||||||||||
| 1B Model, 100B Tokens from Scratch, Training and Evaluation on English + Top 20 Languages | ||||||||||
| FW+FW2 | 0.2743 | 0.2817 | 0.3880 | 0.2586 | 0.5434 | 0.3887 | 0.4086 | 0.3180 | 0.6051 | 0.3852 |
| FWHQ+FW2HQ | 0.2893 | 0.2870 | 0.4461 | 0.2747 | 0.5641 | 0.4035 | 0.4214 | 0.3680 | 0.6308 | 0.4094 |
| FWHQ+FW2HQ + | 0.2952 | 0.2959 | 0.4505 | 0.2816 | 0.5762 | 0.4120 | 0.4116 | 0.3640 | 0.6221 | 0.4121 |
| mFW-edu | 0.2925 | 0.2954 | 0.4766 | 0.2880 | 0.5275 | 0.3999 | 0.4202 | 0.3680 | 0.6022 | 0.4078 |