Model Pretraining

Recent momentum

-38%

31 papers in the last 28 days · 0.5% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

12 new papers

A weekly snapshot of new work published in Model Pretraining.

Period ending 2026-09-14

11 new papers

A weekly snapshot of new work published in Model Pretraining.

Period ending 2026-09-07

13 new papers

A weekly snapshot of new work published in Model Pretraining.

402 papers

Latest in Model Pretraining

  1. A Bitter Lesson for Data Filtering

    May 19, 2026Christopher Mohri, John Duchi, Tatsunori HashimotoLarge Behavior ModelModel Pretraining

  2. LoRA vs. Full Fine-Tuning: A Theoretical Perspective

    May 18, 2026Ali Zindari, Rotem Mulayoff, Sebastian U. StichLow-Rank AdaptationFinetuning

  3. Protein Fold Classification at Scale: Benchmarking and Pretraining

    May 18, 2026Dexiong Chen, Andrei Manolache, Mathias Niepert +1Protein Data BankProtein

  4. TabH2O: A Unified Foundation Model for Tabular Prediction

    May 18, 2026Pascal Pfeiffer, Dmitry Gordeev, Mathias Müller +6Tabular DataModel Pretraining

  5. AMO: Adaptive Muon Orthogonalization

    May 18, 2026Xinlin Zhuang, Panyi Ouyang, Yichen Li +7Muon OptimizerAdam

  6. Scaling Laws for Mixture Pretraining Under Data Constraints

    May 12, 2026Anastasiia Sedova, Skyler Seto, Natalie Schluter +1Model PretrainingScaling Laws

  7. Annotations Mitigate Post-Training Mode Collapse

    May 11, 2026Jacob Mitchell Springer, Madhu Advani, Lukas Aichberger +7LLM Post-Training MethodsModel Pretraining

  8. Hierarchical Mixture-of-Experts with Two-Stage Optimization

    May 8, 2026Gleb Molodtsov, Alexander Miasnikov, Aleksandr BeznosikovMixture-of-Experts ModelsLoad Balancing