Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal Bayesian model leads to near-Bayesian behavior, standard supervised fine-tuning (SFT) on the true answers for the task falls short of it. But behavior alone does not tell us why tuning on a Bayesian or an oracle (true answers) signal differs: whether the resulting LM represents Bayesian beliefs, acts on them, or turns them into a choice the way Bayes' rule does. To compare them, we formulate increasingly demanding requirements for an LM to count as a Bayesian decision maker, spanning its behavior, representations, and computations, and test them on a flight recommendation task. The Bayes-trained LM acts Bayesian, encodes quantities of Bayes' rule in its middle layers, and uses the encoded belief for the recommendation to a certain extent. The oracle-trained LM differs from it both in the beliefs it holds and whether it reads beliefs out into recommendations. Exchanging beliefs between the LMs transfers a part of the Bayesian advantage. Bayes fine-tuning thus installs usable Bayesian beliefs in an LM for reasoning under uncertainty in a way standard SFT on oracle answers cannot, highlighting the advantage of nuanced supervision.
Figures & tables
Figure 1: Overview of our framework for analyzing to which extent LMs instantiate Bayesian inference . We test the framework on a flight recommendation task (left), which is optimally solved through Bayesian decision making (middle). We operationalize the consistency with Bayes-optimal decision making in LMs through four increasingly fine-grained requirements capturing different layers of Bayesian inference (right), showing that LMs are less consistent with increasing requirements.
Figure 2
Figure 4: The quantities of Bayes’ rule emerge in the middle layers. Layer-wise probing for the prior, the likelihood update, and the posterior, each scored against the BayesAssist . Dashed lines mark the reference KL between a uniform distribution and the assistant; the grey area marks the layers 15 – 25 with the main effect; the bands show the deviation across the 20 probes per layer.
Figure 5: Path of the average decoded choice policy ( △ ) of BayesLM and OracleLM and the optimal one under the BayesAssist ( ∙ top corner).
Figure 5
R1
R2
R3
R4
StartLM
×
×
×
×
OracleLM
✓
✓
×
×
BayesLM
✓
✓
(✓)
(✓)
Table 2: Overview of results. Fine-tuning with Bayesian signal, but not oracle one, leads to stronger compliance with the requirements.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
term
β
95% CI of the OR
OR
p
BayesLM
0.8816
[1.8942,3.0781]
2.4147
<10−14
StartLM
−1.4257
[0.1881,0.3070]
0.2403
<10−14
Shuffled
−0.0299
[0.8231,1.1444]
0.9705
1.00
Low-Noise
−0.2102
[0.6620,0.9923]
0.8105
0.209
High-Noise
−0.4215
[0.5239,0.8216]
0.6561
0.002
BayesLM:Shuffled
−0.0062
[0.7899,1.2503]
0.9938
1.00
Appendix
Table 3: Logistic regression analysis of Bayes accuracy in § 3 , regressing Bayes correctness against the LM, the test set and their interaction: correct ∼ C(model) * C(set) . The predictors are dummy coded. OR refers to odds ratio.
term
β
95% CI of the β
OR per SD
p
BayesLM
1.1634
[0.8004,1.5264]
3.2007
<10−14
StartLM
−1.4757
[−1.7296,−1.2219]
0.2286
<10−14
H(policy)
−0.4949
[−0.7101,−0.2798]
0.6096
<10−14
H(policy):BayesLM
−0.6238
[−1.1192,−0.1284]
0.5359
0.0136
H(policy):StartLM
0.4290
[0.1339,0.7241]
1.5357
0.0088
Appendix
Table 4: Logistic regression analysis of policy divergence in § 3 , regressing Bayes correctness against the LM, the entropy of the Bayesian policy and their interaction: correct ∼ C(model) * H(Bayesian policy) . The model is dummy coded. OR refers to odds ratio.
Table 5: Linear regression analysis of Bayes-consistency of decoded beliefs in § 4 , regressing the LM’s token-based behavioral policy against the LM, one of the two predictors (probed policy divergence or prior belief divergence) and their interaction. The model is dummy coded.
Figure 7: Prompt presenting the first round of the recommendation task, marking the probed token positions (reported probing was performed on round 5).
Figure 8: Visualized search over the softmax rationality parameter β for scaling the BayesAssist ’s policy: β=0 is a uniform distribution, β=1 corresponds to the original BayesAssist , β=10 is near-deterministic, and β≈4 is the best match to the LMs’ choice policies.
Figure 9: The BayesLM ’s uncertainty tracks the assistant’s. Entropy of the LM’s behavioral choice policy against the entropy of the Bayesian policy, one point per sample of the original set, colored by whether the recommendation agrees with the BayesAssist . The BayesLM ’s entropy correlates with the assistant’s ( r=0.61 ), the OracleLM ’s only weakly ( r=0.30 ). For both models, the differing instances are concentrated where the BayesAssist itself is most uncertain, but spread more for the OracleLM .
Figure 10: Mean decoded choice policy by layer on the simplex (corners ordered by Bayesian policy), split by whether the LM’s recommendation agrees with the BayesAssist , i.e., is Bayes-correct, (green) or differs (red). Triangles mark the mean decoded choice policy for each layer, and circles mark the mean Bayesian policy of the BayesAssist for the same instances.
Figure 11: Per-layer probing across the four evaluation sets (detail for Figure 4 ). Each row is an evaluation set (original, shuffled, low-noise, high-noise); columns first present probed quantities of Bayes’ rule, and then the probed policy and resulting prediction (argmax of probed policy). On the original and shuffled sets, the belief and choice KL curves fall in the mid-stack and the BayesLM outperforms OracleLM and StartLM in the late layers. On the noisy sets the prior/posterior KL instead rises in the upper layers: the probes are trained on the original set and learn to represent sharper beliefs in late layers, which diverges from the flatter beliefs induced by the noisy sets. Therefore, the KL grows even as choice top-1 stays high — likely a calibration gap of the probe, not a representational failure of the model.
Figure 12: Per-layer probing across token positions (detail for Figure 4 ). Each row is a read-out position (feedback token, last prompt token, digit-generation token); columns first present probed quantities of Bayes’ rule, and then the probed policy and resulting prediction (argmax of probed policy). The choice representation (flight-choice KL, top-1) develops in the mid-to-late layers at the last-prompt and digit tokens — where the BayesLM and OracleLM models rise well above StartLM — but stays near chance at the feedback token, showing the choice is not yet committed there. Belief targets (prior/posterior KL) are decodable across all positions and depths.
Figure 13: Prior probing per feature. KL of the prior probe per layer at the last-input token, on the original set. Dashed: a uniform probe for that feature; grey band: layers 15–25.
setting
flip accuracy
stay accuracy
balanced accuracy
noise effect
Varied dose α , at 32 tokens and 8 layers
α=0.25
0.439
0.953
0.696
0.037
α=0.50
0.817
0.847
0.832
0.037
α=0.75
0.939
0.800
0.870
0.037
α=1.00
0.963
0.765
0.864
0.037
Varied dose α , at 4 tokens and 8 layers
Appendix
Table 6: Hyperparameter search for prototype patching. We select the best hyperparameter for patching using an iterative search, varying one parameter at a time while keeping the rest fixed. We used a separate set of 100 instances where the recommendation should change and 100 where it should remain the same, generated by the BayesAssist . We report for every setting, accuracy on instances that should flip and those that should stay, as well as the effect of injecting a norm-matched random vector (noise effect), We use 32 tokens, α=0.5 and 8 layers. Based on these results, we choose to patch with α=0.5 on 32 tokens across 8 layers, which offers the best trade-off.
Figure 14: Intervening on the encoded belief changes the recommendation in BayesLM . Balanced accuracy results of prototype belief patching.
Figure 15: Prototype patching versus CAA steering. Balanced accuracy of the BayesLM by layer window on the original test set, averaged over the eight feature-direction settings. Prototype patching uses 32 tokens and α=0.5 ; additive steering uses α=0.25 at the 16 tokens after the feedback token or at the digit token. The dotted line marks a model that never changes recommendations; the dashed lines mark the layers where § 4 decodes the belief.
neuro-symbolic
Δ over LM
Llama-3.1-8B
BayesLM
0.86
0.03
OracleLM
0.82
0.10
Qwen-2.5-7B
BayesLM
0.80
0.07
OracleLM
0.74
0.06
Appendix
Table 7: Neuro-symbolic predictions. We read each model’s prior beliefs from its hidden states with linear probes, turn the four decoded distributions into per-attribute rewards, and pick the flight the BayesAssist ’s own rule would pick from them. neuro-symbolic is how often that flight matches the BayesAssist ’s recommendation, at the layer where each model scores best. Δ over LM is the gain over the same model answering on its own.
BayesLM
feature
direction
original
shuffled
low noise
high noise
OracleLM
BayesLM → OracleLM
BayesLM → StartLM
steering
costs
negative
0.775
0.799
0.817
0.818
0.566
0.582
–
0.549
costs
positive
0.769
0.765
0.766
0.766
0.530
0.612
–
0.589
departure time
negative
0.760
0.763
0.751
0.782
0.558
0.662
–
0.565
departure time
positive
0.756
0.773
0.787
0.773
0.533
0.590
–
0.628
duration
negative
0.766
0.770
0.785
0.786
0.505
0.606
0.525
0.613
Appendix
Table 8: Balanced accuracy per feature and direction. Prototype patching at layers 18 – 25 with 32 tokens and α=0.5 . The first four columns patch the BayesLM with its own prior on each test set. The next three use the original test set and patch the OracleLM with its own prior, and the BayesLM ’s prior into the OracleLM and the StartLM . The last column is CAA steering at the 16 tokens after the feedback token ( α=0.25 ) on the original test set. The StartLM recipient was only run for duration and stops.
Figure 16: The choice policy on the simplex, with corners labeled by the BayesAssist ’s ranking of the three options under the updated belief4′ , split by whether the BayesAssist ’s recommendation changes after the intervention. Open circles are the BayesLM ’s mean policy before the intervention, filled circles after the intervention. The diamond marks the BayesAssist ’s own policy before (open) and after (filled).
Figure 17: Overview of the attention masks applied for investigating how posterior beliefs might be formed.
Figure 18: Overview of the attention masking accuracies (y-axis) of the BayesLM and OracleLM , across conversational rounds (x-axis).
Figure 19: Distribution of the choice policy of the BayesLM under masking.
Figure 20: Results of neuro-symbolic predictions across layers. The development of the average Bayes accuracy of the three neuro-symbolic models when the probes stem from different model layers.
Figure 21: Bayes accuracy across backbones (R1, § 3 ). Bayes accuracy of the BayesLM and OracleLM on the four test sets, with 95% confidence intervals. Chance is 1/3 (dashed line). The BayesLM is more Bayes-accurate than the OracleLM on every backbone, and is invariant to evidence order, and behaves as expected under noise.
Figure 22: Oracle accuracy across backbones (R1, § 3 ). Oracle accuracy shows the proportion of recommendations matching the user’s true preference. All models lose accuracy in the presence of noise, as the BayesAssist does, because the evidence is less informative of the user’s preference. The BayesLM exceeds the OracleLM on the original and shuffled sets for every backbone.
Figure 23: Decoded prior across backbones (R2, § 4 ). Belief divergence between the decoded and the BayesAssist ’s prior by layer, on the original set. The dashed line is the divergence of a uniform prior. On every backbone, the prior becomes decodable in the middle layers and more strongly for the BayesLM than for the OracleLM .
Figure 24: Prior patching across backbones (R3, § 5 ). Balanced accuracy of the BayesLM upon intervening on the prior, by the layer window of the edit. The dashed line is a model that never changes its recommendation. The marker shows the best window. The effect is largest for Gemma-2-9B, present for Llama-3.1-8B, and near chance for Qwen-2.5-7B.
Figure 25: Attention masking across backbones (R3, § 5 ). Bayes accuracy without masking and under the one-swoop and sequential masks. Chance is 1/3 (dashed line). On every backbone, both masks reduce accuracy, and the sequential mask reduces it more, though the gap between the two masks is smaller for Llama-3.1-8B and Qwen-2.5-7B than for Gemma-2-9B.
Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration. In this work, we introduce Bayesian-LoRA, which reformulates the deterministic LoRA update as a probabilistic low-rank representation inspired by Sparse Gaussian Processes. We identify a structural isomorphism between LoRA's factorization and Kronecker-factored SGP posteriors, and show that LoRA emerges as a limiting case when posterior uncertainty collapses. We conduct extensive experiments on various LLM architectures across commonsense reasoning benchmarks. With only approximately 0.42M additional parameters and ≈1.2× training cost relative to standard LoRA, Bayesian-LoRA significantly improves calibration across models up to 30B, achieving up to 84% ECE reduction and 76% NLL reduction while maintaining competitive accuracy for both in-distribution and out-of-distribution (OoD) evaluations.
Moule Lin, Shuhao Guan, Andrea Patane +2
School of Computer Science and Statistics, Trinity College Dublin, Dublin, Ireland and Lero the Research Ireland Centre for Software, Ireland · School of Computer Science, University College Dublin, Dublin, Ireland
Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environment. Acting rationally then requires inferring the unobserved quantities that govern it and updating beliefs about them as evidence accumulates. Yet most evaluations only score the model's final-turn answer in a single-turn format, leaving this process unexamined. We ask how closely LLMs' belief updates match those of a rational Bayesian reasoner in multi-turn settings, and introduce BayesBench, a suite of simulation environments that probe this across three progressively complex tasks: (i) Bayesian estimation, where the model infers an unknown parameter from sequential evidence; (ii) Bayesian prediction, where the model turns inferred beliefs about a latent variable into outcome forecasts; and (iii) latent-framed Bayesian prediction, where observations are filtered through a user-persona framing, requiring joint inference over the latent state and the persona. Across seven LLMs (3B--70B), scaling improves latent inference and evidence accumulation, with updates occasionally matching the Bayesian posterior. However, these gains do not reliably carry over to downstream prediction, exposing a gap between inferring latent structure and using it to rationally update beliefs about the target outcome.
Ankur Samanta, Akshayaa Magesh, Tal Lancewicki +7
Meta AI · Columbia University · Work done at Meta +2
Language models are increasingly used in settings where outputs must satisfy user-specified randomness constraints, yet their generation probabilities are often poorly calibrated to those targets. We study whether this capability can be improved directly through fine-tuning. Concretely, we fine-tune language models on synthetic prompts that require sampling from mathematical distributions, and compare two Calibration Fine-Tuning variants: a soft-target method that converts the desired output distribution into trie-derived next-token targets, and a hard-target method that trains on sampled completions from the same target distribution. Across 12 models spanning four families, both methods substantially improve structured-sampling fidelity on held-out distribution families and unseen parameter settings, showing that probabilistic calibration is a trainable capability. Under our selected training configurations, the two methods exhibit different empirical profiles: hard-target fine-tuning is often strongest on structured numeric sampling, while soft-target fine-tuning performs better on broader stochastic generation benchmarks, including open-ended random generation, multiple-choice answer-position balancing, and NoveltyBench. The gains sometimes reduce downstream capability, especially arithmetic reasoning, with costs varying by model. Overall, our results show that probabilistic calibration can be improved through fine-tuning, with our hard-target configuration favoring exact numeric fidelity and our soft-target configuration favoring broader stochastic transfer. Code is available at https://github.com/chandar-lab/calibration-finetuning.