LLM Training

LLM: Large Language Model

Momentum

31 papers in the last four weeks, up 72% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 226

All topics
CardsList
  1. Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling

    May 30, 2026Qiao Xiao, Boqian Wu, Patrik Okanovic +6Language Model PretrainingEfficient Language Model Training

  2. GNMR: Runtime Stability Control for Low-Precision Large Language Model Training

    May 30, 2026Boao Kong, Weichen Jia, Engao Zhang +6LLM Training

  3. Mellum2 Technical Report

    May 29, 2026Marko Kojic, Ivan Bondyrev, Aral de Moor +6Mixture-of-Experts Language ModelsOpen-Weight Language Models

  4. D3^3: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training

    May 29, 2026Yuanjian Xu, Jianing Hao, Guang Zhang +1Efficient Language Model TrainingLLM Training

  5. Unlocking the Working Memory of Large Language Models for Latent Reasoning

    May 28, 2026Lukas Aichberger, Sepp HochreiterMemory-Augmented Language ModelsEfficient Language Model Reasoning

  6. Demystifying Data Organization for Enhanced LLM Training

    May 28, 2026Yalun Dai, Yangyu Huang, Tongshen Yang +8Language Model PretrainingTraining Data Curation

  7. MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

    May 28, 2026Haowen Wang, Yaxin Du, Jian Yang +9Data SelectionEfficient Language Model Training

  8. AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training

    May 28, 2026Ling Chen, Houming Wu, Wenjie YuPipeline ParallelismLLM Training

  9. Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

    May 28, 2026Jing Huang, Daniel Wurgaft, Rachit Bansal +6Long-Tail LearningLanguage Model Scaling Laws

  10. Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

    May 26, 2026Mingze Wang, Shuchen Zhu, Yuxin Fang +3Language Model PretrainingWeight Decay

  11. Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories

    May 26, 2026Sil Hamilton, David MimnoLLM AlignmentDiverse Text Generation

  12. Merge-Bench: Resolve Merge Conflicts with Large Language Models

    May 25, 2026Benedikt Schesch, Michael D. ErnstSoftware EngineeringCode Language Models

  13. BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training

    May 25, 2026Zili Zhang, Chengxu Yang, Shenglong Zhang +8Deep Learning OptimizationMultimodal Large Language Models

  14. When Mean CE Fails: Median CE Can Better Track Language Model Quality

    May 23, 2026Hao Guo, Simon Dennis, Rivaan Patil +1LLM EvaluationLLM Training

  15. Unified Data Selection for LLM Reasoning

    May 21, 2026Xiaoyuan Li, Yubo Ma, Chengpeng Li +6Training Data SelectionLLM Training

  16. One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs

    May 21, 2026Di He, Songjun Tu, Keyu Wang +2Learning Rate SchedulingEfficient Language Model Training

  17. Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

    May 20, 2026Dayal Singh Kalra, Maissam BarkeshliLanguage Model Scaling LawsHyperparameter Transfer

  18. DEL: Digit Entropy Loss for Numerical Learning of Large Language Models

    May 19, 2026Zhaohui Zheng, Chenhang He, Shihao Wang +3Numerical Reasoning in Language ModelsNeural Network Optimization

  19. A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training

    May 19, 2026Daniel Barley, Jonathan Leis, Benjamin Klenk +1Communication-Efficient Distributed TrainingPipeline Parallelism

  20. A Bitter Lesson for Data Filtering

    May 19, 2026Christopher Mohri, John Duchi, Tatsunori HashimotoLanguage Model PretrainingTraining Data Curation

  21. Language models struggle with compartmentalization

    May 19, 2026Thomas Vincent Howe, David WingateMultilingual Language ModelsLLM Interpretability