Information elicitation is an open-ended information-seeking problem in which an interaction can unfold in many potentially valuable directions, requiring an elicitor to continually determine which information to pursue as new information emerges. In agentic elicitation, these decisions may be delegated to a foundation model, yet how model choice shapes the resulting information-seeking behavior remains understudied. We study how judgments about information value vary across LLMs and how these differences shape sequential information seeking. We first examine these judgments across 11 LLMs spanning multiple model families and parameter scales, using a shared set of information and elicitation objectives. We then develop a controlled elicitation simulation in which different models encounter the same information space and use the same selection rule, isolating these judgments from question generation and respondent behavior. Using this setting, we characterize the breadth-depth behavior that emerges from model-specific information-seeking preferences over the course of elicitation. We further examine how interaction history changes the evaluation and subsequent selection of prospective information. We test the robustness and boundaries of these findings through sensitivity analyses and ablations over the opportunities available to the elicitor, the response labels used to operationalize information-seeking preferences, the presence of interaction history, and whether redundancy is explicitly relevant to the assessment. The project code, data, and trajectory files are available at https://github.com/infosenselab/open-elicitation.
Figures & tables
Figure 1: Illustration of the experimental elicitation setup and example trajectories.
Evaluation Topic ↓ Model →
SS
L3
L8
L70
Q7
Q14
Q72
M8
M14
M24
P4
G12
…adopt a multi-party system (original)
.38
.97
.94
1.00
1.00
1.00
1.00
.86
.84
.99
.93
1.00
Positions 2-4
…oppose collectivism
.02
.97
.82
1.00
.00
.00
.00
.72
.75
.84
.69
.99
…introduce compulsory voting
.19
.97
.86
1.00
.11
.00
.16
.75
.75
.68
.73
.84
…adopt libertarianism
.01
.96
.79
1.00
.00
.00
.00
.74
.74
.94
.82
1.00
Positions 35-37
Table 1: Information-seeking scores ( P(Yes) ) across models to “ A multi-party system could lead to more situations where a majority is not reached. This means that it could be harder to push policies through. ” Objectives are shown at the top, middle, and bottom of M14’s ranking; SS denotes semantic similarity. All topics begin with “we should.”
SS
L3
L8
L70
Q7
Q14
Q72
M8
M14
M24
P4
G12
SS
1
0.23
0.16
0.29
0.25
0.28
0.21
0.24
0.32
0.32
0.14
0.34
L3
1
.25
.53
.33
.37
.31
.41
.53
.43
.14
.39
L8
1
.27
.20
.32
.22
.26
.29
.26
.22
.24
L70
1
.37
.39
.47
.50
.60
.51
.13
.49
Q7
1
.28
.33
.31
.36
.34
.18
.38
Q14
1
.37
.34
.47
.45
.22
.42
Table 2: Agreement between models in their information-seeking preferences across objectives, with STSB semantic similarity (SS) included as a reference.
Sel. WA
Cand. WA
Max Rate
Switch
Topics
Run
Rand
.792
.791
.228
.423
32.9
2.34
L3
.856
.792
.339
.348
27.2
2.83
L8
.860
.793
.329
.458
36.5
2.16
L70
.799
.789
.231
.504
39.6
1.96
Q7
.847
.791
.315
.551
43.8
1.79
Q14
.840
.791
.292
.560
44.7
1.76
Table 3: Breadth–depth behavior at k=7 (three-seed means). Sel./Cand.: selected/candidate WA; Max Rate: maximum-WA selection rate; Switch: topic-switch rate; Topics: distinct topics; Run: mean run length.
Metric →
Selected WA
Candidate WA
Candidate Max WA
Selects Max WA
Distinct Topics
k ↓
L8
Q14
M8
L8
Q14
M8
L8
Q14
M8
L8
Q14
M8
L8
Q14
M8
3
.838
.825
.824
.792
.791
.791
.936
.935
.935
.472
.442
.438
29.7
35.4
26.7
5
.853
.832
.831
.793
.790
.789
.970
.970
.969
.369
.327
.325
34.4
41.2
31.2
7
.860
.840
.837
.793
.791
.790
.984
.984
.983
.329
.292
.290
36.5
44.7
34.0
9
.866
.843
.841
.792
.791
.789
.991
.991
.991
.323
.279
.275
37.8
46.0
34.3
11
.868
.846
.843
.793
.792
.790
.995
.995
.995
.312
.271
.265
38.9
47.1
35.3
Table 4: Effect of candidate-set size k on elicitation trajectories. Values are means across seeds 42, 43, and 44; the selected setting ( k=7 ) is underlined . Model abbreviations follow Table 1 . Standard errors are at most .0045 for WA-based measures and .464 topics for Distinct Topics.
W/ History
W/out History
Model
Group
Conf.
Not Conf.
Conf.
Not Conf.
Q7
≤0.7
.471
.485
.670
.667
>0.7
.961
.945
.952
.946
Q14
≤0.7
.430
.481
.622
.619
>0.7
.257
.649
.823
.833
M8
≤0.7
.663
.677
.598
.594
Table 5: Scores by history, semantic similarity, and prior confirmation (three-seed means). Conf.: confirmed; Not Conf.: unconfirmed. Bold indicates contrasts discussed in text.
W/ History
W/out History
Model
StC
SOtU
StC
SOtU
L3
.315
.276
.330
.301
L8
.201
.216
.251
.238
L70
.033
.064
.216
.211
Q7
.287
.233
.280
.265
Q14
.042
.120
.211
.224
Table 6: Selection rates for candidates similar to confirmed (StC) or only to unconfirmed (SOtU) information, with and without history (three-seed means).
Figure 2: Response-label formulation robustness. Mean within-set rank correlation between Yes and Valuable .
W/ History
W/out History
Model
Prompt
Conf.
Not Conf.
Conf.
Not Conf.
Q7
Original
.961
.945
.952
.946
Neutral
.979
.968
.957
.958
Q14
Original
.257
.649
.823
.833
Neutral
.887
.929
.899
.925
M8
Original
.736
.739
.631
.635
Table 7: Scores for highly similar candidates ( >0.7 ) under the original and neutral prompts (three-seed means).
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
W/ Uncert.
W/out Uncert.
Yes
No
Uncert.
Yes
No
Yes Δ
We should adopt a multi-party system
0.94
0.05
0.02
0.93
0.08
0.01
We should introduce compulsory voting
0.86
0.13
0.02
0.79
0.21
0.07
We should oppose collectivism
0.82
0.14
0.04
0.75
0.25
0.07
We should adopt libertarianism
0.79
0.15
0.06
0.75
0.25
0.05
We should adopt an austerity regime
0.71
0.25
0.04
0.50
0.50
0.21
Appendix
Table B.1: Effect of including Uncertain on information-seeking scores. YesΔ denotes the change in P(Yes) when Uncertain is included.
Topic Switch Rate
Distinct Topics
Topic Run Length
Model
Yes
Valuable
∣Δ∣
Yes
Valuable
∣Δ∣
Yes
Valuable
∣Δ∣
L3
.348
.331
0.017
27.177
25.977
1.200
2.834
2.986
0.152
L8
.458
.561
0.103
36.523
44.907
8.384
2.158
1.760
0.398
Q7
.551
.555
0.004
43.810
43.767
0.043
1.794
1.781
0.013
Q14
.560
.521
0.039
44.720
41.680
3.040
1.764
1.897
0.133
M8
.413
.666
0.253
33.997
54.230
20.233
2.542
1.485
1.057
Appendix
Table C.2: Breadth–depth behavior under the original Yes / No / Uncertain labels (Yes) and alternative Valuable / Not valuable / Uncertain labels (Valuable) (three-seed means).
Candidate
History Argument
Sim. Score
The vow of celibacy should be abandoned because it’s outdated in today’s society.
The vow of celibacy is outdated in today’s society and so it should be abandoned.
0.97
We should not ban telemarketing because it would put people out of jobs.
We should not ban telemarketing because it would result in a loss of jobs.
0.97
Burning the flag is freedom of speech and should be allowed.
Flag burning is a form of free speech and should not be banned.
0.96
We should abolish the right to keep and bear arms because there are those who abuse that power and go on rampages killing innocent people and shooting up schools.
We should abolish the right to keep and bear arms because too many abuse that right by killing other and going on rampages and school shootings.
0.95
If we legalise the organ trade then it will mean that more lives can be saved.
Legalizing organ trade has the potential to save many lives.
0.95
People should be allowed to work for as long as they want and as long as they are capable.
Individuals should be able to work as long as they are willing and able.
0.95
Appendix
Table D.3: Ten highest-similarity pairs from Qwen2.5-14B-Instruct trajectories.
Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of "in the wild" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our findings show that LLM responses fail more often in four criteria: sycophantic behavior, overreliance, a default US-centric perspective, and poor handling of vulnerable populations -- with failure rates being mostly higher for analytical queries. By providing methods to monitor the reliability, safety, and fairness of LLM behavior, our dataset and evaluation framework offer an empirical foundation for the broader question of how these systems should behave as they take on a growing role in information access.
Tanise Ceron, Joachim Baumann, Elisa Bassignana +3
1Bocconi University · 2Stanford University · 3IT University of Copenhagen +1
Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across the multi-turn interactions required in sequential decision-making settings. We introduce Amortised Sequential Information Gathering (ASIG), a fine-tuning approach that amortises Bayesian Experimental Design (BED) into LLM policies via a multi-turn extension of Group Relative Policy Optimisation with an Expected Information Gain reward. Evaluated on the 20 Questions task, ASIG more than doubles the success rate of the 7B base model and reduces inference cost by over 25× relative to BED-LLM, a competitive inference-time baseline. Applied to MediQ, a medical diagnosis benchmark unseen during training, ASIG improves information-seeking performance at the 7B scale, suggesting that the learned strategies can transfer out of distribution. Our findings show that amortising BED into LLM policies provides an effective and computationally efficient approach to sequential information gathering.
Jakob Hartmann, James Harvey, Jhonathan Navott +5
University of Oxford, Oxford, United Kingdom · 2Ellison Institute of Technology, Oxford, United Kingdom
Large language models are increasingly used as surrogate models for low-data optimization, but their optimizer-facing prediction and its uncertainty remain poorly understood. We study the surrogate belief elicited from an LLM under sparse observations, showing that it depends strongly on prompt text and query protocol. We introduce an uncertainty-alignment criterion that measures whether model uncertainty tracks residual ambiguity among sample-consistent functions. Across controlled inference tasks and Bayesian optimization studies, we find that structural prompts act as effective priors, POINTWISE and JOINT querying induce different beliefs, and sequential evidence leads to non-monotonic, order-sensitive confidence updates. These effects change downstream acquisition decisions and regret, showing that elicitation protocol is part of the LLM surrogate specification, not a formatting detail.
Ge Lei, Samuel J. Cooper
Dyson School of Design Engineering Imperial College London Exhibition Road, London SW7 2AZ