Rational Clarification by Assistive Agents via Value-of-Information Reasoning
Authors: T. Duy Nguyen-Hien, Yee Whye Teh, Wee Sun Lee, Tan Zhi-Xuan
Organizations: Department of Computer Science, National University of Singapore · Department of Statistics, University of Oxford · Agency for Science, Technology and Research (A*STAR)
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user's intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks --- ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) --- we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.
Figures & tables
Figure 1: Overview of REVOIR. (a) Given an ambiguous user request, REVOIR infers and updates a belief over the user’s intent θ . (b) REVOIR computes the value-of-information (VoI) of asking (vs. acting) by simulating the expected benefit of acting after a clarifying question (vs. a user correction). (c) REVOIR decides between asking or acting by maximizing the cost-adjusted value of asking Vask vs. acting Vact , rationally adapting to cases where: (i) the assistant’s action is terminal; (ii) the user can give corrections.
Figure 2: CondAmbigQA results for Llama-3.1-70B and GPT-5.4-mini . ReAct variants and non-interactive base/top lines are evaluated at five reasoning-effort (R/E) levels for GPT-5.4-mini. Horizontal lines correspond to baselines and toplines across R/E levels: gray lines represent the Direct Answer baselines, and black lines represent Oracle-ReAct toplines, with line styles indicating effort levels: (−⋅⋅−) for none , (\mbox−−−\mbox−−−) for low , (−⋅−) for medium , (⋅⋅⋅) for high , and (−−) for xhigh . REVOIR and REIGN use budgets of (Bagent,Buser)=(100,50) , and SC-BoN-ReAct uses K=5 samples.
Figure 3: Effort-adjusted answer correctness on CondAmbigQA for Llama-3.1-8B as a function of the effort-to-correctness ratio α . The y -axis plots AnswerCorrectness−α⋅Effort for the best method configuration at each value of α . Effort is the number of clarifications per conversation. REVOIR dominates other methods for most values of α under both (a) agent termination ( α>0.008 ) and (b) user termination ( α>0.002 ).
Figure 4: Clarification effort on CondAmbigQA for Llama-3.1-8B across termination variants and methods. When moving from agent termination to user termination , only REVOIR adapts to using less clarifications in total by relying on user corrections.
Figure 5: ADAPT preference satisfaction rate and question count across methods (4-fold cross-validated following ADAPT splits in Patel et al. 2025 ). The standard deviation is computed across split means. Baseline results are reproduced from Table 1 in Patel et al. (2025) . For ICL variants, all seen personas are provided in context during belief updating. More details are in Table I.1 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Split
Examples
Unique Questions
Train
3,062
1,600
Dev
380
200
Test
380
200
Appendix
Table C.1: CondAmbigQA Data Splits
Figure D.1: CondAmbigQA results for Llama-3.1-8B under agent termination. (a) Answer correctness vs. average number of questions asked. Each numbered point corresponds to an agent/user word budget (Bagent,Buser) or (for ReAct) sample budget K ; higher numbers map to higher values; Table D.1 lists exact values. (b) Mean pairwise differences in answer correctness between REVOIR and each baseline across budget points. ReAct is a single operating point broadcast across the budget axis. Shaded bands are bootstrapped standard errors, with Nbootstrap=10,000 .
Figure D.2: CondAmbigQA results for Llama-3.1-8B under user termination. (a) Answer correctness vs. number of user clarifications (questions asked by the agent or corrections issued by the user). Each numbered point corresponds to an agent/user word budget (Bagent,Buser) or (for ReAct) sample budget K ; higher numbers map to higher values; Table D.2 lists exact values. (b) The operating points of budget-aware methods (REVOIR and REIGN) plotted against total word budget difference (words used minus configured budget); negative values indicate the agent acted before exceeding its budget for cost-free words.
Method
Input Tokens
Output Tokens
Total Tokens
Answer Correctness
REVOIR (100, 50)
7.0K
3.5K
10.6K
0.5371
REVOIR (150, 75)
9.6K
4.5K
14.2K
0.5400
REVOIR (200, 100)
11.8K
5.1K
16.8K
0.5435
REVOIR (250, 125)
12.7K
5.5K
18.2K
0.5509
REVOIR (300, 150)
14.8K
6.1K
20.9K
0.5502
REVOIR (350, 175)
15.3K
5.9K
21.3K
0.5448
Appendix
Table D.3: Token usage and answer correctness for REVOIR and SC-BoN-ReAct on Llama-3.1-8B under agent termination. SC-BoN-ReAct’s token cost grows linearly with N , reaching ∼74 K tokens at N=20 — over 3× the cost of the most expensive REVOIR operating point — yet correctness gains are marginal and non-monotonic. REVOIR at comparable token budgets achieves similar or greater correctness, demonstrating that principled ask-vs-act decisions are more token-efficient than inference-time scaling within a fixed policy.
Figure D.3: CondAmbigQA results with Claude Haiku 4.5 and Gemini 3.5 Flash-Lite . ReAct variants and non-interactive base/top lines at available (Gemini 3.5 Flash-Lite) and emulated (Claude Haiku 4.5) reasoning-effort (R/E) levels. Horizontal lines in the top panels correspond to baselines and toplines across R/E levels: gray lines represent the Direct Answer baselines, and black lines represent Oracle-ReAct toplines, with line styles indicating effort levels: (−⋅⋅−) for none (or minimal ), (\mbox−−−\mbox−−−) for low , (−⋅−) for medium , (⋅⋅⋅) for high , and (−−) for xhigh .
Figure D.4: Average word counts of non-interactive one-shot queries from frontier models across reasoning efforts.
Figure D.5: CondAmbigQA results with an inattentive user simulation ( pdismissive=0.2 ) for Llama-3.1-8B and Llama-3.1-70B.
Figure D.6: Llama-3.1-8B CondAmbigQA results in inattentive user simulations with pdismissive∈{0.3,0.4,0.5,0.6,0.7,0.8,0.9} .
Simulated user
Human verdict
Satisfied
Not satisfied
Total
Human: satisfied
64
10
74
Human: not satisfied
7
8
15
Total
71
18
89
Appendix
Table E.1: Human and simulated-user satisfaction verdicts.
Agree
Disagree
Total
Human preference vs. LLM judge ranking
41
16
57
Appendix
Table E.2: Agreement between human answer-quality preferences and AnswerCorrectness .
Configuration
Agent term.
User term.
gpt-5.4-mini
0.85
0.65
Llama3-70B
0.25
0.85
Llama3-8B
0.65
0.45
Appendix
Table G.1: The normalized concentration threshold τ of 1−H(b^t)/Hmax , above which Entropy Thresholding answers immediately, is tuned on the Dev split, per agent model and termination variant.
Configuration
Agent term.
User term.
gpt-5.4-m , K=5
0.9
0.6
Llama3-70B , K=5
0.6
0.7
Llama3-8B , K=5
0.6
0.6
Llama3-8B , K=10
0.5
0.6
Llama3-8B , K=15
0.6
0.5
Llama3-8B , K=20
0.7
0.7
Appendix
Table G.2: Self-consistency threshold τ for SC-BoN-ReAct, tuned on the Dev split.
Test Personas
Method
PSR (%)
#Q
Privileged / Forced
Seen Personas
Always-Ask LLM
52.1 ± 0.9
22.2 ± 0.1
Teacher LLM
65.5 ± 1.1
0.0
Unseen Personas
Always-Ask LLM
50.8 ± 1.6
22.4 ± 0.2
Teacher LLM
65.5 ± 1.9
0.0
Fine-tuned
Appendix
Table I.1: Main results on the ADAPT benchmark (4-fold cross-validation; std across split means). Patel et al. (2025) numbers are reproduced from their Table 1. In-context examples from seen personas injected at hypothesis particle regeneration.