Transformer Layers

Recent momentum

+175%

11 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

6 new papers

A weekly snapshot of new work published in Transformer Layers.

Period ending 2026-09-07

2 new papers

A weekly snapshot of new work published in Transformer Layers.

80 papers

Latest in Transformer Layers

  1. Transformer Heads Looking for Order

    Sep 22, 2026Jasper van Doornmalen, Alexander Kozachinskiy, Corinna Mathwieser +4Transformer Layers

  2. MoRE: Mixture of Reused Experts

    Sep 16, 2026Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace +4Mixture-of-Experts ModelsExperts

  3. LLM Layers Immediately Correct Each Other

    Sep 7, 2026Arjun Patrawala, Jiahai Feng, Erik Jones +1Transformer LayersModel Activations

  4. TACTICL: Task-Aware Compression of Tabular ICL Models

    Aug 11, 2026Mykhailo Koshil, Matthias Feurer, Katharina EggenspergerTabular Foundation ModelsContext Compression

  5. MACRO: Markov Chain Routing of Transformer Layers

    Aug 6, 2026Paweł Batorski, Abtin Pourhadi, Akylgali Aitaza +2Large Language Model RoutingTransformer Layers

  6. Learning through Internalization

    Jun 18, 2026Nikolaos Tsilivis, Nirmit Joshi, Marko Medvedev +2InternalizationTransformer Layers

  7. Parallel Recursive LSTM

    May 16, 2026Tristan Gaudreault, Yongyi MaoParallelRecurrent Neural Networks