Model Activations

Momentum

13 papers in the last four weeks, up 30% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 171

All topics
CardsList
  1. MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge

    Oct 6, 2026Mehmet Emre Akbulut, Johannes Geier, Ulf SchlichtmannModel Activations

  2. Latent space bias directions in LLMs capture confidence, not fairness

    Oct 6, 2026Stephanie Buttigieg, Maeve Madigan, Parameswaran Kamalaruban +1Linear Activation SteeringLarge Language Model Bias

  3. Steering by Influence: Curvature Aware Data Weighting for Activation Steering

    Oct 5, 2026James A. E. Dixon, Stephen J. Roberts, Francesco QuinzanLinear Activation SteeringModel Activations

  4. Latent Information Sharing for Accelerating Federated Learning

    Oct 1, 2026Seungjun Lee, Ensieh Khazaei, Dimitrios Hatzinakos +2Federated LearningModel Activations

  5. Kernelized Activation Steering

    Oct 1, 2026Laziz U. Abdullaev, Minh-Hieu Pham, Bach Do +2Linear Activation SteeringModel Activations

  6. Alignment via Training Against Probes Without Losing Monitorability

    Sep 29, 2026Lena Libon, Alexander Panfilov, Ben Rank +3Safety AlignmentModel Activations

  7. Selecting The Most Informative Tokens in Natural Language Autoencoders

    Sep 29, 2026Federico Torrielli, Gianluca Barmina, Andrea Blasi Núñez +4Explainable AI MethodsModel Activations

  8. LLMs Learn to Evade Latent Monitors from Prior Feedback Alone

    Sep 29, 2026Hugo Lyons Keenan, Christopher Leckie, Sarah ErfaniModel ActivationsEvasion

  9. Persona Dosing: Calibrated Activation Steering for Graded Trait Control

    Sep 28, 2026Zehao Jin, Junran Wang, Ruixuan Deng +4PersonalityLinear Activation Steering

  10. Signatures of semantic search in the activations of large language models

    Sep 28, 2026Luke Leckie, Peter M. Todd, Jacob G. FosterModel ActivationsFluency

  11. From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features

    Sep 28, 2026Dewen Liu, Zixuan Li, Jonathan Pan +4Mechanistic InterpretabilityModel Activations

  12. Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

    Sep 27, 2026Gert Lek, Zixuan Xia, Pin-Yu Chen +1Model ActivationsExplainable AI Methods

  13. Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

    Sep 16, 2026Devesh Tiwari, Camille Davis, Shivank Sinha +3Model ActivationsCausal Reasoning

  14. What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

    Sep 14, 2026Lingheng Du, Yiming Tang, Xufeng Duan +1Large Language Model Reinforcement LearningImproving Sparse Autoencoders

  15. GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

    Sep 11, 2026Xuan Cuong Ngo, Hao Vo, Ngan LeLinear Activation SteeringModel Activations

  16. SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery

    Sep 4, 2026Mansooreh Montazerin, Antonio Ortega, Ajitesh SrivastavaSymbolic RegressionModel Discovery

  17. GAPS: Dimension-Level Gates for Conditional Activation Steering

    Sep 1, 2026Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad +1Linear Activation SteeringModel Activations

  18. Interpretable Symptom Vectors for Depression in a Large Language Model

    Sep 1, 2026Fangyi Zhu, Ajay Subramanian, Allison Constant +3DepressionModel Activations

  19. Investigating Assistant Bias in LLM User Simulators Using a Role Vector

    Sep 1, 2026Daeheon Jeong, Yoonjoo Lee, Eugene Choi +2User SimulationBiases

  20. Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

    Aug 31, 2026Ashwin Nedungadi, Stefan Oehmcke, Stefan LüdtkeImage-To-Code GenerationSpatial Reasoning

  21. Probing and steering biology across Boltz-1s trunk-diffusion boundary

    Aug 11, 2026Piotr Jedryszek, Tongmeng Xie, Adam Winnifrith +5Protein Structure PredictionModel Activations

  22. Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

    Aug 11, 2026Nikolai Bolik, Lennart Stöpler, Artur AndrzejakSemantic RepresentationsModel Activations

  23. Interpreting Language Model Hidden States at Scale

    Aug 10, 2026Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson +3Model ActivationsHidden States

  24. A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

    Aug 10, 2026Leandro de Souza Rosa, Lorenzo Capelli, Clara Nunes Barrancos +2Final Convolutional LayersOut-Of-Distribution Detection

  25. One Adapter Pair per Model: A Universal Activation Interface for Language Models

    Aug 10, 2026Su-Hyeon Kim, Jiwan Mun, Yo-Sub HanModel Activations

  26. Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models

    Aug 9, 2026Muhammad Faishal Adly Nelwan, Alfan Farizki WicaksonoLinear Activation SteeringModel Activations

  27. Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes

    Aug 7, 2026Luc Hazenoot, Zhaochun Ren, Amirhossein ZohrehvandModel ActivationsText Analysis and Detection

  28. Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

    Aug 5, 2026Kotaro Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto +3Introspection AdaptersModel Fine-Tuning

  29. Calliphony: A Calligraphy-Driven Interface for Real-Time Generative Music Performance

    Aug 4, 2026Tristan Wu, Ruiji Yu, Gus XiaGenerative MusicHandwriting

  30. Rewriting or Reweighting? A Geometric Account in Language Models

    Aug 3, 2026Juntong Wang, Shengkun Yang, Xiyuan Wang +1Behavioral Foundation ModelsModel Activations

  31. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

    Jul 30, 2026Pere Martra, Eugenio Martínez Cámara, Alfonso Ureña LópezLarge Language Model BiasModel Activations

  32. Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs

    Jul 30, 2026Jinyi Liu, Wei Chen, Pengyu Chen +4LLM Inference OptimizationModel Activations

  33. Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

    Jul 28, 2026Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal +1Model ActivationsLarge Language Model Safety

  34. Forecasting Side Effects of Activation Steering

    Jul 28, 2026Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang +1Linear Activation SteeringSteering

  35. Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

    Jul 27, 2026Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty +2Sparse Autoencoder FeaturesInterpretability

  36. Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

    Jul 23, 2026Renuka Oladri, Niveda Jawahar, Abdirisak MohamedChain-of-Thought ReasoningDeepseek

  37. Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

    Jul 22, 2026Seonglae Cho, Zekun Wu, Kleyton Da Costa +3Sparse Autoencoder FeaturesModel Activations

  38. Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence

    Jul 20, 2026Katarzyna Filus, Sebastian PokucińskiMechanistic InterpretabilitySemantic Representations

  39. Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent

    Jul 17, 2026Sriram Balasubramanian, Soheil FeiziInterpretabilityModel Activations

  40. Transcoders for Investigating Deception in Language Models

    Jul 16, 2026Darius Lim, Nathan Leow, Xin Wei ChiaDeceptionMechanistic Interpretability

  41. Conditional Optimal Bridge for Riemannian Activation Steering

    Jul 12, 2026Seyed Arshan Dalili, Ajay Narayanan Sridhar, Vijaykrishnan Narayanan +1Linear Activation SteeringModel Activations

  42. Training, Reading, and Editing Legible Transformers

    Jul 9, 2026Mark OskinTransformer ArchitecturesModel Activations

  43. Prompt Compression via Activation Aggregation

    Jul 9, 2026Thibaud Ardoin, Semira Einsele, Evis Bregu +1Large Language Model CompressionInstruction-Tuned Models

  44. A First-Principles Theory of Slow Thinking and Active Perception

    Jul 9, 2026Hongkang Yang, Zhi-Qin John Xu, Feiyu Xiong +1Cognitive ScienceModel Activations

  45. Distributed Sparse Interventions in Language Models

    Jul 8, 2026Maximilian S. Ernst, Lorenz Linhardt, Aaron Peikert +1Latent Feature InterventionsNeurons

  46. Unsupervised Features Mining via Activation Geometry

    Jul 5, 2026Amit LeVi, Elad David, Max FominModel ActivationsLLM Reasoning Strategies

  47. kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail

    Jul 2, 2026Mahmoud Abdelfattah, Hamid Nasiri, Peter GarraghanLarge Language Model SafetyAdversarial Prompts

  48. Mechanistically Eliciting Latent Behaviors in Language Models

    Jun 28, 2026Andrew Mack, Nina Panickssery, Alexander Matt TurnerModel ActivationsAgent Behavior Modeling