Language Model Probing

Latest papers 192

All topics
CardsList
  1. PROMPT2BOX: Uncovering Entailment Structure among LLM Prompts

    Mar 22, 2026Neeladri Bhuiya, Shib Sankar Dasgupta, Andrew McCallum +1LLM EvaluationRepresentation Learning

  2. World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models

    Mar 4, 2026Elan BarenholtzLLM World ModelsWord Embeddings

  3. Probing for Knowledge Attribution in Large Language Models

    Feb 26, 2026Ivo Brink, Alexander Boer, Dennis UlmerLLM InterpretabilityHallucination in Language Models

  4. Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry

    Jan 30, 2026Zhuochun Li, Yong Zhang, Ming Li +8LLM EvaluationLLM-as-a-Judge

  5. Training-free Truthfulness Detection via Sparse MLP Value Vectors

    Sep 22, 2025Runheng Liu, Heyan Huang, Xingchen Xiao +2LLM ReliabilityLLM Hallucination Detection

  6. Human Psychometric Questionnaires Mischaracterize LLM Behavior

    Sep 12, 2025Woojung Song, Dongmin Choi, Yoonah Park +3Personality Modeling in Language ModelsPsychometric Validation

  7. The Trilemma of Truth in Large Language Models

    Jun 30, 2025Germans Savcisens, Tina Eliassi-RadLLM InterpretabilityLLM Reliability

  8. Evaluating LLMs on Chinese Topic Constructions: A Research Proposal Inspired by Tian et al. (2024)

    Apr 21, 2025Xiaodong YangLLM EvaluationLanguage Model Probing

  9. Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

    Oct 3, 2024Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva +5Self-AttentionMultiple-Choice Question Answering