Large Models

Recent momentum

+31%

17 papers in the last 28 days · 0.3% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

8 new papers

A weekly snapshot of new work published in Large Models.

Period ending 2026-09-14

2 new papers

A weekly snapshot of new work published in Large Models.

Period ending 2026-09-07

8 new papers

A weekly snapshot of new work published in Large Models.

192 papers

Latest in Large Models

Open your feed →
CardsList
  1. The Life of a Token: from Words to Bits on the Wire

    Sep 17, 2026Davide Avesani, Pengwenlong Gu, Sotiris Skaperas +1Large ModelsNeural Architectures

  2. MiST: Mid-Training LLMs for Cybersecurity

    Sep 16, 2026Oded Ovadia, Elad Ben Zaken, Elad Guttman +1Security AnalysisLarge Models

  3. OPEN-1B: A Fully Auditable Training Run

    Sep 15, 2026John Donaghy, Brian Wilcox, Oğuzhan Ersoy +6ReproducibilityLarge Models

  4. Breaking the 1.58-bit Barrier for Ternary LLMs

    Sep 14, 2026Evangelos Georganas, Alexander Heinecke, Pradeep DubeyLarge Language Model QuantizationLarge Models

  5. Parallelism Strategy Chaining for Fast Training Convergence

    Sep 7, 2026Minchul Kang, Changyong Shin, Younghun Go +4Large ModelsParallel

  6. Instella-MoE Technical Report

    Sep 1, 2026Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra +10Mixture-Of-Expert Language ModelsLarge Models

  7. Interpreting Language Model Hidden States at Scale

    Aug 10, 2026Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson +3Large ModelsModel Activations

  8. A Defense of the Quadratic Model

    Jul 23, 2026Alexandru Meterez, Pranav Ajit Nair, Depen Morwani +3Large ModelsCondition Number

  9. Test-Time Scaling via Error Localization

    Jul 23, 2026Rajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta +2Test-Time ScalingInference-Time

  10. xHC: Expanded Hyper-Connections

    Jul 16, 2026Xiangdong Zhang, Xiaohan Qin, Sunan Zou +10Transformer Residual StreamsResidual Stream

  11. Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models

    Jul 14, 2026Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi +9Large ModelsParallel

  12. Understanding Layer Patching in Model Size Interpolation

    Jul 9, 2026Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak +3Large Models

  13. QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling

    Jul 1, 2026Michael Y. Li, Anthony Zhan, Kanishk Gandhi +2Monte CarloTest-Time Scaling

  14. Smooth Scaling Laws Hide Stepwise Token Learning

    Jun 29, 2026Pingjie Wang, Zechen Hu, Peiru Yang +2Scaling LawsLarge Models

  15. On Test-Time Scaling for Vision-Language Models

    Jun 27, 2026Fawaz Sammani, Tzoulio Chamiti, Nikos DeligiannisTest-Time ScalingLarge Vision Language Models

  16. Can Scale Save Us From Plasticity Loss in Large Language Models?

    Jun 23, 2026J. Fernando Hernandez-Garcia, Tomás Figliolia, Beren MillidgeLarge ModelsPlasticity