Language Model Calibration

Latest papers 247

All topics
CardsList
  1. TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models

    Oct 8, 2026Rahul Sharma, Andrew B. Ducan, Gaétan Marceau Caron +1LLM Decision-MakingLanguage Model Calibration

  2. RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

    Oct 8, 2026Gukhyeon Lee, SangKeun LeeLanguage Model CalibrationLLM Uncertainty Estimation

  3. System Switch: When Should a Fast Decision Model Stop and Think?

    Oct 7, 2026Gian Luca BailoLLM Decision-MakingSelective Prediction

  4. Language-model ratings of depression reflect the rater more than the patient

    Oct 6, 2026Baihan LinInter-Rater ReliabilityAnnotator Disagreement

  5. JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications

    Oct 5, 2026Steven Denney, Matthew DiGiuseppeLLM EvaluationComputational Social Science

  6. Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing

    Oct 5, 2026Andrea Paganelli, Stefano Civelli, Pietro Bernardelle +1Confidence Estimation in Language ModelsLanguage Model Calibration

  7. Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

    Oct 1, 2026Ionel Eduard Stan, Paolo NapoletanoLanguage Model CalibrationComputational Argumentation

  8. HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix

    Oct 1, 2026Zirui He, Haiyan Zhao, Jingyu Hu +5Language Model SteeringLanguage Model Calibration

  9. AnyJev Technical Report

    Sep 30, 2026Jiamu Zhang, Tianze Yang, Yucheng Shi +6Multiple-Choice Question AnsweringLanguage Model Calibration

  10. Evaluating Persistent Calibration under Evolving Model Knowledge

    Sep 30, 2026Victor Wang, Thomas Hofweber, Mohit Bansal +1Language Model Calibration

  11. Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

    Sep 29, 2026Feiyang Li, Shengjing Liu, Qi Zhan +7Uncertainty QuantificationConfidence Estimation in Language Models

  12. Evaluating and Benchmarking the System One Model Jev

    Sep 29, 2026Tobias Deußer, Lorenz Sparrenberg, Rafet SifaMultilingual Language Model EvaluationLLM Evaluation

  13. Persona Dosing: Calibrated Activation Steering for Graded Trait Control

    Sep 28, 2026Zehao Jin, Junran Wang, Ruixuan Deng +4Language Model SteeringPersonality Modeling in Language Models

  14. IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing

    Sep 28, 2026Yung-Chin Chen, Chia-Yu Chen, Naveen VermaLLM QuantizationCompute-in-Memory

  15. TRACE: Single-Pass Decoding-Trace Risk Localization for Generation Calibration

    Sep 28, 2026Yuebin Xu, Xuemei Peng, Junlan Chen +2Confidence Estimation in Language ModelsLanguage Model Calibration

  16. Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models

    Sep 28, 2026Lucas Biechy, Cédric Eichler, Adrien Boiret +1Language Model CalibrationLLM Uncertainty Estimation

  17. Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts

    Sep 28, 2026Chenxiao Fan, Chongming Gao, Gangyi Zhang +8Confidence Estimation in Language ModelsRL for Language Model Reasoning

  18. LLMs learn different forms of metacognition when trained to predict their own accuracy

    Sep 27, 2026Nicolas Yax, Stefano Palminteri, Pierre-Yves OudeyerConfidence Estimation in Language ModelsLanguage Model Calibration

  19. Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs

    Sep 27, 2026Shiyu Ni, Keping Bi, Jiafeng Guo +4Confidence Estimation in Language ModelsLanguage Model Calibration

  20. Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models

    Sep 27, 2026Yadong Xi, Rongsheng Zhang, Tangjie Lv +2Confidence Estimation in Language ModelsLanguage Model Calibration

  21. CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels

    Sep 24, 2026Xiangwei Wang, Peng Wang, Saman HalgamugeLanguage Model Calibration

  22. Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

    Sep 24, 2026Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh +2Prompt SensitivityLLM-as-a-Judge