Language Model Probing

Latest papers 192

All topics
CardsList
  1. Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

    May 27, 2026Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta +4LLM InterpretabilityAdversarial Attacks on LLMs

  2. Can LLMs Introspect? A Reality Check

    May 25, 2026Shashwat Singh, Tal Linzen, Shauli RavfogelLLM EvaluationLLM Interpretability

  3. Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams

    May 25, 2026Tianda Sun, Dimitar KazakovLLM Tool UseLanguage Model Probing

  4. MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing

    May 24, 2026Riasad Alvi, Nurul Labib Sayeedi, Md. Faiyaz Abdullah SayeediHallucination DetectionHallucination in Language Models

  5. Relational Linear Properties in Language Models: An Empirical Investigation

    May 21, 2026Giovanni Valer, Luigi Gresele, Marco Bronzini +1LLM InterpretabilityFactual Knowledge in Language Models

  6. Represented Is Not Computed: A Causal Test of Candidate Algorithmic Intermediates in a Transformer

    May 21, 2026Ishita Darade, Sushrut ThoratNumerical Reasoning in Language ModelsTransformer Interpretability

  7. Do LLMs Know What Luxembourgish Borrows? Probing Lexical Neology in Low-Resource Multilingual Models

    May 20, 2026Nina Hosseini-KivananiMultilingual Language Model EvaluationKnowledge Augmentation for Language Models

  8. Reading Calibrated Uncertainty from Language Model Trajectories

    May 19, 2026Aliai Eusebi, Alexander Herzog, Xiaoyu Liang +3Confidence Estimation in Language ModelsLLM Reliability

  9. Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

    May 18, 2026Maciej Chrabąszcz, Aleksander Szymczyk, Marcin Sendera +2Language Model Safety EvaluationLLM Interpretability

  10. Probing for Representation Manifolds in Superposition

    May 18, 2026Alexander ModellFeature SuperpositionNeural Representation Geometry

  11. Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry

    May 18, 2026Woo Seob Sim, Yu Rang ParkLanguage Model Safety EvaluationLLM Safety

  12. Polar probe linearly decodes semantic structures from LLMs

    May 13, 2026Pablo J. Diego-Simón, Pierre Orhan, Emmanuel Chemla +2Representation LearningLinear Probing

  13. Probing Persona-Dependent Preferences in Language Models

    May 13, 2026Oscar Gilg, Pierre Beckmann, Daniel Paleka +1LLM AlignmentLLM Interpretability

  14. A Controlled Counterexample to Strong Proxy-Based Explanations of OOD Performance: in a Fixed Pretraining-and-Probing Setup

    May 12, 2026Hongmin LiOOD GeneralizationLanguage Model Probing

  15. Instructions Shape Production of Language, not Processing

    May 11, 2026Andreas Waldis, Leshem Choshen, Yufang Hou +1LLM InterpretabilityInstruction Following

  16. Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities

    May 11, 2026Daniel RanardLLM EvaluationNext-Token Prediction

  17. Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings

    May 11, 2026Benjamin Icard, Lila Sainero, Alice Breton +2LLM EvaluationText Embeddings

  18. Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance

    May 11, 2026Shiqiang Wang, Herbert Woisetschläger, Hans Arno Jacobsen +1Synthetic Data GenerationLLM Training

  19. Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions

    May 11, 2026Andrew Lee, Fernanda Viégas, Martin WattenbergTransformer InterpretabilityNeural Representation Geometry

  20. Do Linear Probes Generalize Better in Persona Coordinates?

    May 10, 2026Prasad Mahadik, Adrians SkaparsLanguage Model Safety EvaluationDeception in Language Models

  21. The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

    May 9, 2026Rania Elbadry, Ahmed Heakl, Fan Zhang +4LLM InterpretabilityLLM Reliability

  22. Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

    May 9, 2026Nils A. Herrmann, Leander Girrbach, Kirill Bykov +1Language Model SteeringLLM Interpretability