cs.AISep 29, 2026

Diagnosing and Improving Probabilistic Reasoning in Large Language Models

Authors: Huaman Sun, Dingcheng Wang, Jason Hartline, Jessica Hullman

Organizations: Department of Computer Science Northwestern University Evanston, IL 60201, USA

Abstract

Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL interventions targeting beliefs, decisions, or both improve these components across three domains, whether improvements transfer across components and elicitation formats, and whether decision performance can improve without improvement in belief formation. We find that targeting one component of probabilistic reasoning redistributes decision loss, improving the target without necessarily transferring to others, and that jointly targeting belief formation and decision-making improves both but hinges on matched formats between training and evaluation.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How reliable are LLMs when it comes to playing dice?

    Jun 5, 2026Luca Avena, Gianmarco Bet, Bernardo BusoniLLM Reasoning StrategiesProbabilistic Model

  2. Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

    May 7, 2026Ömer Faruk Akgül, Rajgopal Kannan, Willie Neiswanger +1LLM Reasoning StrategiesReasoning Skills

  3. LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs

    May 7, 2026Chacha Chen, Matthew Jörke, Adam Goliński +4LLM Reasoning StrategiesBayesian