Language Model Introspection

Latest papers 22

All topics
CardsList
  1. Identifying Introspection From the Inside

    Oct 5, 2026David I. Atkinson, Dillon Plunkett, David BauLLM InterpretabilityLanguage Model Introspection

  2. A mechanistic study of language model introspection

    Sep 28, 2026Jiahong Zou, Xiangkun Sun, Lingkai Kong +1Attention Head AnalysisLLM Interpretability

  3. Evaluating and Improving LLM Self-Modeling

    Aug 31, 2026Siqi Zeng, Andre N. Assis, Rowan WangLLM EvaluationLanguage Model Introspection

  4. "Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

    Aug 8, 2026Adelaide Danilov, Aria Nourbakhsh, Oleksandr Marchenko Breneur +1Large Language Model-Based Role-Play SimulationPersonality Modeling in Language Models

  5. Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

    Aug 5, 2026Kotaro Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto +3LLM AuditingLanguage Model Introspection

  6. Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement

    Jul 5, 2026Jiang Zhang, Bing Yuan, Qian ZhangLanguage Model IntrospectionLanguage Model Self-Improvement

  7. Revealing Hidden Model Behaviors with Task-Specific Self-Reports

    Jul 3, 2026Taras Kutsyk, Bartosz ZielińskiLLM AuditingLanguage Model Introspection

  8. ICA Lens: Interpreting Language Models Without Training Another Dictionary

    Jun 10, 2026Sida Liu, Feijiang HanIndependent Component AnalysisLLM Interpretability

  9. Can LLMs Introspect? A Reality Check

    May 25, 2026Shashwat Singh, Tal Linzen, Shauli RavfogelLLM EvaluationLLM Interpretability

  10. Tool Calling is Linearly Readable and Steerable in Language Models

    May 8, 2026Zekun Wu, Ze Wang, Seonglae Cho +4AI Agent ReliabilityTool-Augmented Language Model Agents

  11. TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models

    Apr 30, 2026Amirreza Esmaeili, Fatemeh FardLLM InterpretabilityCode Generation

  12. Introspection Adapters: Training LLMs to Report Their Learned Behaviors

    Apr 18, 2026Keshav Shenoy, Li Yang, Abhay Sheshadri +4LLM AuditingLanguage Model Introspection