Transformer Architectures

Recent momentum

-24%

52 papers in the last 28 days · 0.8% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

20 new papers

A weekly snapshot of new work published in Transformer Architectures.

Period ending 2026-09-14

1 new paper

A weekly snapshot of new work published in Transformer Architectures.

Period ending 2026-09-07

2 new papers

A weekly snapshot of new work published in Transformer Architectures.

684 papers

Latest in Transformer Architectures

  1. Pretraining Curricula Enable Selective Fine-tuning

    Jul 6, 2026Sebastian A. Bruijns, Jirko Rubruck, Mia H. Whitefield +3CurriculumBalanced Learning

  2. ELiTeFormer: An Efficient Transformer for FPGAs

    Jul 4, 2026Victor Agostinelli, Nicolas Bohm Agostini, Antonino TumeoField-Programmable Gate ArraysLLM Inference Efficiency

  3. Induction Heads Interpolate N-Grams

    Jul 2, 2026Francesco D'Angelo, Oguz Kaan Yuksel, Swathi Shree Narashiman +1Transformer ArchitecturesIn-Context Learning

  4. The State-Prediction Separation Hypothesis

    Jul 1, 2026Giovanni Monea, Nathan Godey, Kianté Brantley +1Transformer ArchitecturesNext Activity Prediction

  5. Hierarchical Global Attention (HGA)

    Jun 29, 2026Woernle Frank, Fedosov Vladimir, Grinenko ArtemiyEfficient Long-Context InferenceTransformer Architectures