cs.LGJul 3, 2026

Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs

Authors: Jakob HartmannJames HarveyJhonathan NavottErik Y. WangLuckeciano C. MeloFlaviu CipciganCheng ZhangAlessandro Abate

Organizations: University of Oxford, Oxford, United Kingdom · 2Ellison Institute of Technology, Oxford, United Kingdom

Abstract

Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across the multi-turn interactions required in sequential decision-making settings. We introduce Amortised Sequential Information Gathering (ASIG), a fine-tuning approach that amortises Bayesian Experimental Design (BED) into LLM policies via a multi-turn extension of Group Relative Policy Optimisation with an Expected Information Gain reward. Evaluated on the 20 Questions task, ASIG more than doubles the success rate of the 7B base model and reduces inference cost by over 25×25\times relative to BED-LLM, a competitive inference-time baseline. Applied to MediQ, a medical diagnosis benchmark unseen during training, ASIG improves information-seeking performance at the 7B scale, suggesting that the learned strategies can transfer out of distribution. Our findings show that amortising BED into LLM policies provides an effective and computationally efficient approach to sequential information gathering.

Explore similar work

CardsList
  1. CA-BED: Conversation-Aware Bayesian Experimental Design

    May 31, 2026Daniel Arnould, Rashad Aziz, Zixuan Kang +5Bayesian Experimental Design