Language Model Probing

Latest papers 192

All topics
CardsList
  1. Tool Calling is Linearly Readable and Steerable in Language Models

    May 8, 2026Zekun Wu, Ze Wang, Seonglae Cho +4AI Agent ReliabilityTool-Augmented Language Model Agents

  2. The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models

    May 7, 2026Chonghan Qin, Xiachong Feng, Ziyun Song +3LLM InterpretabilityLanguage Model Probing

  3. When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment

    May 7, 2026Long Zhang, Wei-neng Chen, Feng-feng Wei +1LLM InterpretabilityLanguage Model Probing

  4. HyperLens: Quantifying Cognitive Effort in LLMs with Fine-grained Confidence Trajectory

    May 7, 2026Chengda Lu, Xiaoyu Fan, Wei XuConfidence Estimation in Language ModelsLLM Interpretability

  5. Implicit Representations of Grammaticality in Language Models

    May 6, 2026Yingshan Susan Wang, Linlu Qiu, Zhaofeng Wu +2Language Model Probing

  6. Retrieval and competition: how a protein foundation model starts a protein

    May 5, 2026Piotr Jedryszek, Oliver M. CrookProtein Language ModelsMechanistic Interpretability

  7. Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

    May 1, 2026Gaofei Shen, Martijn Bentum, Tomas O. Lentz +2Transformer InterpretabilityRepresentation Probing

  8. Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives

    Apr 28, 2026Yu Wang, Emmanuele Chersoni, Chu-Ren HuangMultilingual Language Model EvaluationLLM Grounding

  9. The Pragmatic Persona: Discovering LLM Persona through Bridging Inference

    Apr 27, 2026Jisoo Yang, Jongwon Ryu, Minuk Ma +2Personality Modeling in Language ModelsLanguage Model Probing

  10. Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms

    Apr 23, 2026Yuto Nishida, Naoki Shikoda, Yosuke Kishinami +4LLM EvaluationMemorization in Language Models

  11. Measuring Opinion Bias and Sycophancy via LLM-based Persuasion

    Apr 23, 2026Rodrigo Nogueira, Giovana Kerche Bonás, Thales Sales Almeida +7LLM SycophancyLLM Auditing

  12. Slot Machines: How LLMs Keep Track of Multiple Entities

    Apr 22, 2026Paul C. Bogdan, Jack LindseyLLM InterpretabilityLanguage Model Probing

  13. Tracing Relational Knowledge Recall in Large Language Models

    Apr 21, 2026Nicholas Popovič, Michael FärberLinear ProbingFactual Knowledge in Language Models

  14. Probing for Reading Times

    Apr 20, 2026Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re +4Computational PsycholinguisticsLanguage Model Probing

  15. Dual Alignment Between Language Model Layers and Human Sentence Processing

    Apr 20, 2026Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki +1Language ModelingComputational Psycholinguistics

  16. STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

    Apr 20, 2026Sungeun An, Swanand Ravindra Kadhe, Shailja Thakur +2LLM EvaluationReasoning Evaluation

  17. SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

    Apr 20, 2026Mohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Orojlooyjadid +2LLM EvaluationBenchmark Contamination

  18. Introspection Adapters: Training LLMs to Report Their Learned Behaviors

    Apr 18, 2026Keshav Shenoy, Li Yang, Abhay Sheshadri +4LLM AuditingLanguage Model Introspection

  19. Segment-Level Coherence for Robust Harmful Intent Probing in LLMs

    Apr 16, 2026Xuanli He, Bilgehan Sel, Faizan Ali +3LLM SafetyLLM Jailbreak Attacks

  20. What are They Thinking? Delineation, Probing, and Tracking of Concepts in LLMs

    Apr 7, 2026Mohamed Abdelwahab, Michelle Yu Collins, Sihan Chen +5Linear ProbingLLM Interpretability

  21. GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics

    Apr 3, 2026Yujing Wang, Yuanbang Liang, Yukun Lai +2LLM InterpretabilityLLM Uncertainty Estimation

  22. ReLope: From Hidden-State Probing to a Decision Module for Multimodal LLM Routing

    Mar 25, 2026Yaopei Zeng, Congchao Wang, Blake JianHang Chen +1Multimodal Large Language ModelsLLM Routing