Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance. While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data. This work explores extending quality filtering to over 100 languages by proposing a multilingual adaptation approach that converts an existing English quality classifier into a multilingual variant. Our approach proposes training a small multi-layer perceptron on top of Transformer encoder-only model embeddings, using multilingual text as input and scores obtained from English classifiers applied to machine-translated text as labels. Our 1B, 3B and 8B scale experiments show that our approach maintains the downstream LLM benchmark performance of existing multilingual model-based filtering baselines, without harming regional and cultural knowledge benchmarks. To further evaluate cross-lingual generalization, we compare classifier scores of high-quality synthetic data and web samples, and the correlation of classifier scores with LLM-based ones, revealing that the classifier can learn the scoring criteria of its original English variant, even for languages not included in its training data.
Figures & tables
GMMLU c
INCLUDE c
ARC
mARC
HS
mHS
XNLI
OBQA
XWG
Avg.
Classifier
Emb. Model
1B Model, 100B Tokens from Scratch, Training and Evaluation on English + Top 20 Languages
mFW-edu
XLM-R
0.2814
0.2944
0.4812
0.2936
0.5412
0.4034
0.4161
0.3800
0.6046
0.4106
mmBERT
0.2989
0.2913
0.4897
0.2937
0.5469
0.4026
0.4121
0.3740
0.6180
0.4141
mDCLM
XLM-R
0.2845
0.2965
0.4802
0.2815
0.5651
0.4080
0.4227
0.3640
0.6277
0.4145
mmBERT
0.2848
0.2938
0.4554
0.2744
0.5454
0.4012
0.4142
0.3780
0.6314
0.4087
Table 1: Benchmark performance of 1B parameter multilingual LLMs trained on 100B tokens. We compare models trained on data selected using our multilingually adapted mFW-edu and mDCLM quality classifiers, paired with XLM-RoBERTa and mmBERT embeddings. We retain the top 10% of documents based on the classifier scores.
Figure 1: Benchmark performance of 1B parameter multilingual LLMs trained on 100B tokens. We compare models trained on data selected using our multilingually adapted mFW-edu classifier with mmBERT embeddings trained on a varying amount of languages, and the FW+FW2, FWHQ+FW2HQ, and FWHQ+FW2HQ + baselines. We retain the top 10% of documents based on the classifier scores, except for FW+FW2 which is not model-filtered.
GMMLU c
INCLUDE c
ARC
mARC
HS
mHS
XNLI
OBQA
XWG
Avg.
Model
1B Model, 100B Tokens from Scratch, Training and Evaluation on English + Top 20 Languages
FW+FW2
0.2743
0.2817
0.3880
0.2586
0.5434
0.3887
0.4086
0.3180
0.6051
0.3852
FWHQ+FW2HQ
0.2893
0.2870
0.4461
0.2747
0.5641
0.4035
0.4214
0.3680
0.6308
0.4094
FWHQ+FW2HQ +
0.2952
0.2959
0.4505
0.2816
0.5762
0.4120
0.4116
0.3640
0.6221
0.4121
mFW-edu (Ours)
0.2989
0.2913
0.4897
0.2937
0.5469
0.4026
0.4121
0.3740
0.6180
0.4141
Table 2: Benchmark performance of 1B, 3B, and 8B parameter multilingual LLMs. We compare models trained on data selected using our multilingually adapted mFW-edu and mDCLM quality classifiers paired with mmBERT embeddings, and the FW+FW2, FWHQ+FW2HQ, and FWHQ+FW2HQ + baselines. We retain the top 10% of documents based on the classifier scores, except for FW+FW2 which is not model-filtered.
Chinese
Japanese
French
Portuguese
Avg.
Model
1B Model, 30B Tokens From Scratch, Monolingual Training and Evaluation
FWHQ+FW2HQ +
0.3578
0.3453
0.3938
0.3140
0.3527
mFW-edu (Ours)
0.3358
0.3653
0.4153
0.3230
0.3598
mFW-edu Eng (Ours)
0.3541
0.3453
0.3771
0.3339
0.3526
Table 3: Benchmark performance of 1B monolingual Chinese, Japanese, French and Portuguese LLMs on the INCLUDE benchmark. We compare models trained on data selected using our multilingually adapted mFW-edu quality classifier paired with mmBERT embeddings trained on all data (mFW-edu), trained on English (mFW-edu Eng ), and the FWHQ+FW2HQ + baseline. We retain the top 10% of documents based on the classifier scores.
China
Japan
France
Brazil
Avg. (All Subsets)
Model
8B Model, 63B Token Cooldown from 9.5T Token Checkpoint, Training and
Evaluation on All Available Languages
Pre-cooldown
0.4915
0.6792
0.6429
0.4000
0.5764
FWHQ+FW2HQ +
0.5424
0.7925
0.6429
0.3600
0.6143
mFW-edu (Ours)
0.5763
0.7170
0.5714
0.4000
0.6164
Table 4: Benchmark performance of 8B multilingual LLMs on the easy subset of the CulturalBench benchmark. We compare models trained on data selected using our multilingually adapted mFW-edu quality classifier paired with mmBERT embeddings trained on all data (mFW-edu), trained on English (mFW-edu Eng ), and the FWHQ+FW2HQ + baseline. We retain the top 10% of documents based on the classifier scores.
Figure 2: Comparison of quality scores between our mFW-edu classifier with mmBERT embeddings across selected languages-script pairs (ISO 639-3) among the top 100, evaluated on FineWeb (FW), FineWeb 2 (FW2), and synthetic encyclopedia data.
Figure 3: Comparison of our mFW-edu classifier with mmBERT embeddings against LLM-based scores ( gemma-3-27b-it , gpt-oss-120b , Qwen3.5-35B-A3B ). Top row: in-distribution (top 100), bottom row: out-of-distribution language-script pairs (beyond top 100; ISO 639-3).
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Score distribution of the original classifiers and the multilingual adaptations (left), score error distribution (middle), and the correlation between between their scores (right) on English FineWeb data.
Figure 5: Score distribution of our multilingually adapted mFW-edu classifier with mmBERT embeddings on the English synthetic encyclopedia data.
GMMLU c
INCLUDE c
ARC
mARC
HS
mHS
XNLI
OBQA
XWG
Avg.
Model
1B Model, 100B Tokens from Scratch, Training and Evaluation on English + Top 20 Languages
FW+FW2
0.2743
0.2817
0.3880
0.2586
0.5434
0.3887
0.4086
0.3180
0.6051
0.3852
FWHQ+FW2HQ
0.2893
0.2870
0.4461
0.2747
0.5641
0.4035
0.4214
0.3680
0.6308
0.4094
FWHQ+FW2HQ +
0.2952
0.2959
0.4505
0.2816
0.5762
0.4120
0.4116
0.3640
0.6221
0.4121
mFW-edu Eng
0.2925
0.2954
0.4766
0.2880
0.5275
0.3999
0.4202
0.3680
0.6022
0.4078
Appendix
Table 5: Benchmark performance of 1B parameter multilingual LLMs trained on 100B tokens. We compare models trained on data selected using our multilingually adapted mFW-edu classifier paired with mmBERT embeddings trained on a varying amount of languages, and the FW+FW2, FWHQ+FW2HQ, and FWHQ+FW2HQ + baselines. We retain the top 10% of documents based on the classifier scores, except for FW+FW2 which is not model filtered.
As Large Language Models (LLMs) scale, data curation has shifted from maximizing volume to optimizing the signal-to-noise ratio by performing quality filtering. However, for many languages, native high quality data is insufficient to train robust quality classifiers. This work investigates the idea that quality markers in embedding space may show cross-lingual consistency, which would allow high-resource languages to subsidize the filtering of low-resource ones. We evaluate various filtering strategies, including cross-lingual transfer, third quartile sampling (Q3), and retention rate tuning. Our results demonstrate that massive multilingual pooling frequently outperforms monolingual baselines in both rank stability and aggregate accuracy for a 1B model trained on 103B tokens, delivering gains for high resource languages (1.2% increase in aggregate normalized accuracy for French) and matching or exceeding monolingual baselines for low-resource languages. However, we find that scale alone does not guarantee stability. Furthermore, for high-resource languages like French, we show that refining the decision boundary through third quartile sampling (Q3) or tuning the retention rate is necessary to fully leverage the multilingual signal.
Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to English within the model. Several cross-lingual alignment (CLA) scores have been proposed for use with LLMs, along with multiple approaches for extracting embeddings from the models. We provide a comparative analysis of 27 CLA score variants, examining how they differ and how well each predicts downstream performance across three tasks. Crucially, while LLMs are widely used for generative tasks such as machine translation, prior work has focused almost exclusively on classification. We therefore investigate whether CLA scores are similarly predictive of translation performance. To enable computing correlations across target languages, we propose a PMI-based translation metric, which is less dependent on the target language and correlates strongly with chrF. We find that CLA with English predicts translation quality comparably to or better than source-target CLA, providing new evidence that LLMs use English as an internal pivot language.
Adnan Al Ali, Kathy Hämmerl, Jindřich Libovický +1
Charles University, Faculty of Mathematics and Physics, Czech Republic · Technical University of Munich, Germany · Munich Center for Machine Learning
Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities across languages. We present a systematic, controlled study of the interplay between training language coverage, model scale, and task domain, based on 220 supervised fine-tuning runs on parallel translated multilingual data mixtures spanning mathematical reasoning and API calling tasks, with models up to 8B parameters. We find that English-only post-training is typically suboptimal: incorporating even a single non-English language improves both English performance and cross-lingual generalization. Increasing language diversity during post-training generally yields further gains, particularly for low-resource languages, while performance on high-resource languages tends to plateau rather than degrade. Moreover, greater language diversity enables strong zero-shot transfer to unseen languages, reducing the need for direct inclusion, though gains remain limited for typologically distant, low-resource languages.