28 papers in the last four weeks, up 180% on the four weeks before. 0.3% of all new papers.
Feb 4, 2026·Jonas Rohweder, Subhabrata Dutta, Iryna GurevychTransformer InterpretabilityLatent Variable Models
UKP Lab, TU Darmstadt
Feb 2, 2026·Anthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai +1Transformer InterpretabilityKernel Regression
ECE Department, UT Austin · Department of Physics, UT Austin · Oden Institute, UT Austin
Jan 9, 2026·Nora Graichen, Iria de-Dios-Flores, Gemma BoledaTransformer InterpretabilityLLM Evaluation
Universitat Pompeu Fabra1 · ICREA2
Dec 23, 2025·Sophie ZhaoTransformer InterpretabilityNeural Representation Geometry
School of Computer Science Georgia Institute of Technology
Nov 22, 2025·Adrian Goldwaser, Michael Munn, Javier Gonzalvo +1Transformer InterpretabilityTransformer FFNs
University of Cambridge, UK · Google Research
Oct 21, 2025·Brady Bhalla, Honglu Fan, Nancy Chen +1Transformer InterpretabilityRepresentation Learning
California Institute of Technology · Pasadena, CA 91125 · Google DeepMind +3
Oct 10, 2025·Davide Maltoni, Matteo FerraraTransformer InterpretabilityCoT Reasoning
Department of Computer Science and Engineering, University of Bologna, Italy
Sep 29, 2025·Pengxiao Lin, Zheng-An Chen, Zhi-Qin John XuTransformer InterpretabilityOOD Generalization
School of Mathematical Sciences, Shanghai Jiao Tong University · Institute of Natural Sciences, MOE-LSC, Shanghai Jiao Tong University · Shanghai Seres Information Technology Co., Ltd, Shanghai 200040, China.
Sep 22, 2025·Sara Papi, Dennis Fucci, Marco Gaido +2Transformer InterpretabilityCross-Attention
Fondazione Bruno Kessler, Italy
Jun 10, 2025·Li-Ming Zhan, Bo Liu, Chengqiang Xie +2Transformer InterpretabilityLanguage Model Steering
Department of Data Science and Artificial Intelligence The Hong Kong Polytechnic University Hong Kong S.A.R.
Feb 20, 2025·Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev +4Transformer InterpretabilityLong-Context Language Modeling
AIRI · Skoltech · Lomonosov Moscow State University +1
Oct 31, 2024·Ambroise Odonnat, Wassim Bouaziz, Vivien CabannesTransformer InterpretabilityAttention Mechanisms
Inria, Univ. Rennes 2 · Mistral AI; work done while at Meta. · FAIR at Meta.
Date pending·Yongjin Cui, Xiaohui FanGradient-Based AttributionTransformer Interpretability
Zhejiang University
Date pending·Corentin Kervadec, Iuliia Lysova, Iuri Macocco +2Transformer InterpretabilityLLM Interpretability
Universitat Pompeu Fabra · ICREA