cs.CLSep 30, 2026

Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow

Authors: Polina Tsvilodub, Andreas Waldis, Linlu Qiu, Tal Linzen, Michael Franke

Organizations: University of Tübingen · Massachusetts Institute of Technology · New York University

Abstract

Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal Bayesian\textit{Bayesian} model leads to near-Bayesian behavior, standard supervised fine-tuning (SFT) on the true answers for the task falls short of it. But behavior alone does not tell us why\textit{why} tuning on a Bayesian\textit{Bayesian} or an oracle\textit{oracle} (true answers) signal differs: whether the resulting LM represents Bayesian beliefs, acts on them, or turns them into a choice the way Bayes' rule does. To compare them, we formulate increasingly demanding requirements for an LM to count as a Bayesian decision maker, spanning its behavior, representations, and computations, and test them on a flight recommendation task. The Bayes-trained LM acts Bayesian, encodes quantities of Bayes' rule in its middle layers, and uses the encoded belief for the recommendation to a certain extent. The oracle-trained LM differs from it both in the beliefs it holds and whether it reads beliefs out into recommendations. Exchanging beliefs between the LMs transfers a part of the Bayesian advantage. Bayes fine-tuning thus installs usable Bayesian beliefs in an LM for reasoning under uncertainty in a way standard SFT on oracle answers cannot, highlighting the advantage of nuanced supervision.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jan 28, 2026cs.AI

Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration. In this work, we introduce Bayesian-LoRA, which reformulates the deterministic LoRA update as a probabilistic low-rank representation inspired by Sparse Gaussian Processes. We identify a structural isomorphism between LoRA's factorization and Kronecker-factored SGP posteriors, and show that LoRA emerges as a limiting case when posterior uncertainty collapses. We conduct extensive experiments on various LLM architectures across commonsense reasoning benchmarks. With only approximately 0.42M additional parameters and ≈1.2×{\approx}1.2{\times} training cost relative to standard LoRA, Bayesian-LoRA significantly improves calibration across models up to 30B, achieving up to 84% ECE reduction and 76% NLL reduction while maintaining competitive accuracy for both in-distribution and out-of-distribution (OoD) evaluations.
Jun 29, 2026cs.AI

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environment. Acting rationally then requires inferring the unobserved quantities that govern it and updating beliefs about them as evidence accumulates. Yet most evaluations only score the model's final-turn answer in a single-turn format, leaving this process unexamined. We ask how closely LLMs' belief updates match those of a rational Bayesian reasoner in multi-turn settings, and introduce BayesBench, a suite of simulation environments that probe this across three progressively complex tasks: (i) Bayesian estimation, where the model infers an unknown parameter from sequential evidence; (ii) Bayesian prediction, where the model turns inferred beliefs about a latent variable into outcome forecasts; and (iii) latent-framed Bayesian prediction, where observations are filtered through a user-persona framing, requiring joint inference over the latent state and the persona. Across seven LLMs (3B--70B), scaling improves latent inference and evidence accumulation, with updates occasionally matching the Bayesian posterior. However, these gains do not reliably carry over to downstream prediction, exposing a gap between inferring latent structure and using it to rationally update beliefs about the target outcome.
May 12, 2026cs.CL

Probabilistic Calibration Is a Trainable Capability in Language Models

Language models are increasingly used in settings where outputs must satisfy user-specified randomness constraints, yet their generation probabilities are often poorly calibrated to those targets. We study whether this capability can be improved directly through fine-tuning. Concretely, we fine-tune language models on synthetic prompts that require sampling from mathematical distributions, and compare two Calibration Fine-Tuning variants: a soft-target method that converts the desired output distribution into trie-derived next-token targets, and a hard-target method that trains on sampled completions from the same target distribution. Across 12 models spanning four families, both methods substantially improve structured-sampling fidelity on held-out distribution families and unseen parameter settings, showing that probabilistic calibration is a trainable capability. Under our selected training configurations, the two methods exhibit different empirical profiles: hard-target fine-tuning is often strongest on structured numeric sampling, while soft-target fine-tuning performs better on broader stochastic generation benchmarks, including open-ended random generation, multiple-choice answer-position balancing, and NoveltyBench. The gains sometimes reduce downstream capability, especially arithmetic reasoning, with costs varying by model. Overall, our results show that probabilistic calibration can be improved through fine-tuning, with our hard-target configuration favoring exact numeric fidelity and our soft-target configuration favoring broader stochastic transfer. Code is available at https://github.com/chandar-lab/calibration-finetuning.