Period ending 2026-09-21
3 new papers
A weekly snapshot of new work published in Large Language Model Pretraining.
Twelve weeks of publication activity for this topic as it is defined today.
Weekly history
What was published in this topic, kept on the site without email delivery.
Period ending 2026-09-21
A weekly snapshot of new work published in Large Language Model Pretraining.
Period ending 2026-09-14
A weekly snapshot of new work published in Large Language Model Pretraining.
Period ending 2026-09-07
A weekly snapshot of new work published in Large Language Model Pretraining.
116 papers
learners'' that execute local inner optimization steps. These learners asynchronously communicate parameter fragments to a central synchronizer, which circumvents failed or straggling learners by aggregating updates using a minimum quorum, an adaptive grace window, and dynamic token-weighted merging. Inspired by chaos engineering'', we achieve significantly improved training efficiency in failure-prone environments with millions of simulated chips with strictly zero global downtime, while maintaining competitive model performance across text and vision tasks, for both dense and mixture-of-expert architectures.Vikhr'' refers to the name of the Mistral LLM series and means a strong gust of wind.'' Unlike previous Russian-language models that typically rely on LoRA adapters on top of English-oriented models, sacrificing performance for lower training costs, Vikhr features an adapted tokenizer vocabulary and undergoes the continued pre-training and instruction tuning of all weights. This not only enhances the model's performance but also significantly improves its computational and contextual efficiency. We also expanded the instruction datasets and corpora for continued pre-training. The model weights, instruction sets, and code are publicly available.