Reliable decision-making requires more than accurate prediction: a model must preserve its beliefs, apply the relevant utilities, and recognize when the information needed to justify an action is missing. We study whether language models can learn this decision procedure from supervised fine-tuning and generalize it across domains and differing natural-language expressions of the decision challenge. Across 20 datasets, we explore challenges of belief instability and decision-making errors by first eliciting probabilities of outcomes and then varying only the utilities and the framing of the decision problems, while holding the evidence fixed. We train models to preserve elicited beliefs while selecting the action that maximizes expected utility, and evaluate transfer to unseen application domains, held-out framings, and different classes of payoff structures. We further introduce incomplete-information settings in which required utilities are withheld and replaced with irrelevant text, testing whether models can distinguish missing decision-relevant information from merely additional context. We find that targeted fine-tuning substantially improves coherent decision-making and that in many situations, learning transfers across domains and framings to situations unobserved during training. Further, models trained for decidability learn to identify when action cannot be justified based on missing information. Finally, we show the value of a routed system that considers separately the recognition of decision completeness and utility-sensitive decision execution.
Figures & tables
Figure 1: Overview of the benchmark and training pipeline. (a) We first fine-tune a complete-information solver to preserve its elicited belief and apply the utility-implied decision rule across eight equivalent framings of each case–utility pair. (b) We then study decision decidability under incomplete utility information. The decidability model identifies which information condition is present, reports a belief, withholds action for incomplete cases, and routes complete decision problems to the specialist solver.
Figure 3: Complete-information results for Direct SFT (top row of each panel) and CoT SFT (bottom row), pooled over the eight framing holdouts: (a) belief-conditional decision error (Equation 6 ); (b) excess belief distortion error (Equation 5 ). The three left bars split test prompts by framing status (seen, held-out cross-class, held-out full-class) and the two right bars by domain status. Dashed lines are the same base model without SFT on the same prompts, reweighted to the test set’s domain mix. Error bars are 95% bootstrap intervals.
Figure 4: Decidability results by model size (solid: Gemma 4; dashed: Qwen3.5). Rows: (a) information-choice accuracy over all four conditions, (b) belief-conditional decision error on true-D items, (c) end-to-end joint accuracy, and (d) excess belief distortion error over all items that report a belief, against a floor measured on the decidability prompts. The routed system’s information accuracy is its gate’s (a), its decisions are the complete-information CoT specialist’s (b), and its floor is the route-weighted mixture of its components’ floors (d). Error bars are 95% bootstrap intervals, often smaller than the markers.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Domain
Role sentence opening each prompt
Event E
Training domains (15): train, validation, and in-domain test splits
Credit-card default risk
You are a credit-risk analyst at a bank reviewing a customer’s credit-card account to judge repayment risk.
the cardholder defaults on their next payment
Loan credit risk
You are a loan officer at a retail bank reviewing a new loan application.
the applicant turns out to be a bad credit risk
Corporate bankruptcy
You are a corporate credit analyst assessing whether a company is financially distressed from its accounting ratios.
the company goes bankrupt within the forecast horizon
Term-deposit marketing
You are a marketing analyst at a retail bank judging whether a client will subscribe to a term deposit if called.
the client subscribes to a term deposit if contacted
Income-support targeting
You are an analyst screening census records to identify lower-income individuals for a support program.
the person earns at most $50K a year (low income)
Appendix
Table 1: Decision domains, role sentences, and events. Ten cases are sampled per domain, giving 200 evidence contexts.
Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal Bayesian model leads to near-Bayesian behavior, standard supervised fine-tuning (SFT) on the true answers for the task falls short of it. But behavior alone does not tell us why tuning on a Bayesian or an oracle (true answers) signal differs: whether the resulting LM represents Bayesian beliefs, acts on them, or turns them into a choice the way Bayes' rule does. To compare them, we formulate increasingly demanding requirements for an LM to count as a Bayesian decision maker, spanning its behavior, representations, and computations, and test them on a flight recommendation task. The Bayes-trained LM acts Bayesian, encodes quantities of Bayes' rule in its middle layers, and uses the encoded belief for the recommendation to a certain extent. The oracle-trained LM differs from it both in the beliefs it holds and whether it reads beliefs out into recommendations. Exchanging beliefs between the LMs transfers a part of the Bayesian advantage. Bayes fine-tuning thus installs usable Bayesian beliefs in an LM for reasoning under uncertainty in a way standard SFT on oracle answers cannot, highlighting the advantage of nuanced supervision.
Polina Tsvilodub, Andreas Waldis, Linlu Qiu +2
University of Tübingen · Massachusetts Institute of Technology · New York University
Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL interventions targeting beliefs, decisions, or both improve these components across three domains, whether improvements transfer across components and elicitation formats, and whether decision performance can improve without improvement in belief formation. We find that targeting one component of probabilistic reasoning redistributes decision loss, improving the target without necessarily transferring to others, and that jointly targeting belief formation and decision-making improves both but hinges on matched formats between training and evaluation.
Huaman Sun, Dingcheng Wang, Jason Hartline +1
Department of Computer Science Northwestern University Evanston, IL 60201, USA
Large language models (LLMs) are increasingly used as reasoning modules in many applications. While they are efficient in certain tasks, LLMs often struggle to produce human-aligned solutions. Human-aligned decision making requires accounting for both explicitly stated goals and latent user preferences that shape how ambiguous situations should be resolved. Existing approaches to incorporating such preferences either rely on extensive and repeated user interactions or fail to generalize latent preferences across tasks and contexts, limiting their practical applicability. We consider a setting in which an LLM is used for high-level reasoning and is responsible for inferring latent user preferences from limited interactions, which guides downstream decision making. We introduce CLIPR (Conversational Learning for Inferring Preferences and Reasoning), a framework that learns actionable, transferable natural language rules that represent latent user preferences from minimal conversational input. These rules are iteratively refined through adaptive feedback and applied to both in-distribution and out-of-distribution ambiguous tasks across multiple environments. Evaluations on three datasets and a user study show that CLIPR consistently outperforms existing methods in improving alignment and reducing inference costs.