Language Model Probing

Latest papers 192

All topics
CardsList
  1. Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog

    Oct 8, 2026Hadeel Al-Negheimish, Jasna Ilieva, Yoon KimLLM ReasoningLong-Context Language Model Inference

  2. Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals

    Oct 6, 2026Jerzy Kamiński, Ilya Galyukshev, Artem Kuznetsov +6Multi-Turn Dialogue EvaluationLanguage Model Probing

  3. Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing

    Oct 6, 2026Jing Chen, Giulia Loca, Simona Amenta +1Language ModelingComputational Psycholinguistics

  4. Will the Judge Flip? Predicting Position-Sensitive LLM Judgments from Residual Stream Activations

    Oct 5, 2026Hashmath Shaik, Gnaneswar Villuri, Alex DoboliLLM EvaluationLLM-as-a-Judge

  5. Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing

    Oct 5, 2026Andrea Paganelli, Stefano Civelli, Pietro Bernardelle +1Confidence Estimation in Language ModelsLanguage Model Calibration

  6. External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

    Oct 1, 2026Kingshuk Gupta, Davide BuscaldiHallucination DetectionLLM Hallucination Detection

  7. Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability

    Oct 1, 2026Yu Mao, Lei Yu, Zining Zhu +6RL for Language ModelsReinforcement Learning

  8. Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

    Sep 30, 2026Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek +1Deception in Language ModelsLLM Reliability

  9. Locating Answer-Correctness Signals in Frozen Large Language Models

    Sep 29, 2026Yuansen Liu, Yixuan Tang, Anthony Kum Hoe TungLLM Answer VerificationLanguage Model Error Detection

  10. Reader Proficiency Shapes Layer-wise Surprisal Profiles

    Sep 29, 2026Akio Hayakawa, Horacio SaggionLanguage Model Probing

  11. Rethinking Reasoning Paths as Phase-Structured Trajectories

    Sep 29, 2026Zhenghao He, Guangzhi Xiong, Sanchit Sinha +3Structured ReasoningLanguage Model Probing

  12. Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax

    Sep 24, 2026Zhenyan Lu, He Wang, Xiaohui HuangLLM EvaluationLanguage Model Decoding

  13. Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

    Sep 24, 2026Huseyin Cavus, Sebin Sabu, Joshua Spear +2LLM InterpretabilityHallucination in Language Models

  14. The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts

    Sep 21, 2026Manjiang Yu, Hongji Li, Zihan Wang +5LLM InterpretabilityLinear Representation Hypothesis

  15. Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

    Sep 16, 2026Alizishaan Khatri, Chiquita Prabhu, Omkar NeogiSafety FilteringLLM Safety

  16. From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

    Sep 12, 2026Wenkang Wei, Yuan Fang, Renhe Jiang +2LLM InterpretabilityFactual Knowledge in Language Models

  17. Legible Failures: Detecting and Repairing In-Context Binding Errors

    Sep 11, 2026Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. HariLanguage Model Error DetectionLLM Reliability

  18. Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise

    Sep 7, 2026Mika Okamoto, Gabriele SartiTransformer InterpretabilityLanguage Model Probing

  19. Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA

    Sep 1, 2026Nishant Mishra, Ameen Abu-Hanna, Iacer CalixtoClinical Language Model EvaluationLanguage Model Robustness

  20. Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models

    Sep 1, 2026Tian Fang, Gaël Guibon, Davide BuscaldiCausal Interventions in Language ModelsEmotion Representation in Language Models