Model Activations

Recent momentum

-40%

9 papers in the last 28 days · 0.1% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

2 new papers

A weekly snapshot of new work published in Model Activations.

Period ending 2026-09-14

3 new papers

A weekly snapshot of new work published in Model Activations.

Period ending 2026-09-07

5 new papers

A weekly snapshot of new work published in Model Activations.

158 papers

Latest in Model Activations

  1. Inside the LLM Word Factory

    Jun 7, 2026Benzi Busigin, Yuval PinterTokenizationModel Activations

  2. Steering Vectors are an Adversarial Attack Surface

    Jun 4, 2026Abzal Aidakhmetov, Donato Crisostomi, Tommaso Mencattini +3Language Model Evasion AttacksActivation Steering

  3. Covert Influence Between Language Models

    Jun 2, 2026Avidan Shah, Jay Chooi, Jinghua Ou +1InfluenceTraining Language Models

  4. The role of class encoding in neural collapse

    May 29, 2026Bastien Massion, Roy Makhlouf, Estelle MassartNeural CollapseCollapse

  5. Building Better Activation Oracles

    May 23, 2026Jan Bauer, Celeste De Schamphelaere, Adam Karvonen +2Model ActivationsModel Interpretability

  6. Vision-Language Binding in In-Context Image Generation

    May 23, 2026Chris Ge, Rohit Gandikota, Antonio Torralba +1Single ImagesVisual Context

  7. Where Does Authorship Signal Emerge in Encoder-Based Language Models?

    May 19, 2026Francis Kulumba, Guillaume Vimont, Laurent Romary +1AuthorshipEncoding Models

  8. Tracing Persona Vectors Through LLM Pretraining

    May 13, 2026Viktor Moskvoretskii, Dominik Glandorf, Jorge Medina Moreira +2PersonalityLarge Language Model Safety