Rational Clarification by Assistive Agents via Value-of-Information Reasoning
Authors: T. Duy Nguyen-Hien, Yee Whye Teh, Wee Sun Lee, Tan Zhi-Xuan
Organizations: Department of Computer Science, National University of Singapore · Department of Statistics, University of Oxford · Agency for Science, Technology and Research (A*STAR)
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user's intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks --- ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) --- we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.
Figures & tables
Figure 1: Overview of REVOIR. (a) Given an ambiguous user request, REVOIR infers and updates a belief over the user’s intent θ . (b) REVOIR computes the value-of-information (VoI) of asking (vs. acting) by simulating the expected benefit of acting after a clarifying question (vs. a user correction). (c) REVOIR decides between asking or acting by maximizing the cost-adjusted value of asking Vask vs. acting Vact , rationally adapting to cases where: (i) the assistant’s action is terminal; (ii) the user can give corrections.
Figure 2: CondAmbigQA results for Llama-3.1-70B and GPT-5.4-mini . ReAct variants and non-interactive base/top lines are evaluated at five reasoning-effort (R/E) levels for GPT-5.4-mini. Horizontal lines correspond to baselines and toplines across R/E levels: gray lines represent the Direct Answer baselines, and black lines represent Oracle-ReAct toplines, with line styles indicating effort levels: (−⋅⋅−) for none , (\mbox−−−\mbox−−−) for low , (−⋅−) for medium , (⋅⋅⋅) for high , and (−−) for xhigh . REVOIR and REIGN use budgets of (Bagent,Buser)=(100,50) , and SC-BoN-ReAct uses K=5 samples.
Figure 3: Effort-adjusted answer correctness on CondAmbigQA for Llama-3.1-8B as a function of the effort-to-correctness ratio α . The y -axis plots AnswerCorrectness−α⋅Effort for the best method configuration at each value of α . Effort is the number of clarifications per conversation. REVOIR dominates other methods for most values of α under both (a) agent termination ( α>0.008 ) and (b) user termination ( α>0.002 ).
Figure 4: Clarification effort on CondAmbigQA for Llama-3.1-8B across termination variants and methods. When moving from agent termination to user termination , only REVOIR adapts to using less clarifications in total by relying on user corrections.
Figure 5: ADAPT preference satisfaction rate and question count across methods (4-fold cross-validated following ADAPT splits in Patel et al. 2025 ). The standard deviation is computed across split means. Baseline results are reproduced from Table 1 in Patel et al. (2025) . For ICL variants, all seen personas are provided in context during belief updating. More details are in Table I.1 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Split
Examples
Unique Questions
Train
3,062
1,600
Dev
380
200
Test
380
200
Appendix
Table C.1: CondAmbigQA Data Splits
Figure D.1: CondAmbigQA results for Llama-3.1-8B under agent termination. (a) Answer correctness vs. average number of questions asked. Each numbered point corresponds to an agent/user word budget (Bagent,Buser) or (for ReAct) sample budget K ; higher numbers map to higher values; Table D.1 lists exact values. (b) Mean pairwise differences in answer correctness between REVOIR and each baseline across budget points. ReAct is a single operating point broadcast across the budget axis. Shaded bands are bootstrapped standard errors, with Nbootstrap=10,000 .
Figure D.2: CondAmbigQA results for Llama-3.1-8B under user termination. (a) Answer correctness vs. number of user clarifications (questions asked by the agent or corrections issued by the user). Each numbered point corresponds to an agent/user word budget (Bagent,Buser) or (for ReAct) sample budget K ; higher numbers map to higher values; Table D.2 lists exact values. (b) The operating points of budget-aware methods (REVOIR and REIGN) plotted against total word budget difference (words used minus configured budget); negative values indicate the agent acted before exceeding its budget for cost-free words.
Method
Input Tokens
Output Tokens
Total Tokens
Answer Correctness
REVOIR (100, 50)
7.0K
3.5K
10.6K
0.5371
REVOIR (150, 75)
9.6K
4.5K
14.2K
0.5400
REVOIR (200, 100)
11.8K
5.1K
16.8K
0.5435
REVOIR (250, 125)
12.7K
5.5K
18.2K
0.5509
REVOIR (300, 150)
14.8K
6.1K
20.9K
0.5502
REVOIR (350, 175)
15.3K
5.9K
21.3K
0.5448
Appendix
Table D.3: Token usage and answer correctness for REVOIR and SC-BoN-ReAct on Llama-3.1-8B under agent termination. SC-BoN-ReAct’s token cost grows linearly with N , reaching ∼74 K tokens at N=20 — over 3× the cost of the most expensive REVOIR operating point — yet correctness gains are marginal and non-monotonic. REVOIR at comparable token budgets achieves similar or greater correctness, demonstrating that principled ask-vs-act decisions are more token-efficient than inference-time scaling within a fixed policy.
Figure D.3: CondAmbigQA results with Claude Haiku 4.5 and Gemini 3.5 Flash-Lite . ReAct variants and non-interactive base/top lines at available (Gemini 3.5 Flash-Lite) and emulated (Claude Haiku 4.5) reasoning-effort (R/E) levels. Horizontal lines in the top panels correspond to baselines and toplines across R/E levels: gray lines represent the Direct Answer baselines, and black lines represent Oracle-ReAct toplines, with line styles indicating effort levels: (−⋅⋅−) for none (or minimal ), (\mbox−−−\mbox−−−) for low , (−⋅−) for medium , (⋅⋅⋅) for high , and (−−) for xhigh .
Figure D.4: Average word counts of non-interactive one-shot queries from frontier models across reasoning efforts.
Figure D.5: CondAmbigQA results with an inattentive user simulation ( pdismissive=0.2 ) for Llama-3.1-8B and Llama-3.1-70B.
Figure D.6: Llama-3.1-8B CondAmbigQA results in inattentive user simulations with pdismissive∈{0.3,0.4,0.5,0.6,0.7,0.8,0.9} .
Simulated user
Human verdict
Satisfied
Not satisfied
Total
Human: satisfied
64
10
74
Human: not satisfied
7
8
15
Total
71
18
89
Appendix
Table E.1: Human and simulated-user satisfaction verdicts.
Agree
Disagree
Total
Human preference vs. LLM judge ranking
41
16
57
Appendix
Table E.2: Agreement between human answer-quality preferences and AnswerCorrectness .
Configuration
Agent term.
User term.
gpt-5.4-mini
0.85
0.65
Llama3-70B
0.25
0.85
Llama3-8B
0.65
0.45
Appendix
Table G.1: The normalized concentration threshold τ of 1−H(b^t)/Hmax , above which Entropy Thresholding answers immediately, is tuned on the Dev split, per agent model and termination variant.
Configuration
Agent term.
User term.
gpt-5.4-m , K=5
0.9
0.6
Llama3-70B , K=5
0.6
0.7
Llama3-8B , K=5
0.6
0.6
Llama3-8B , K=10
0.5
0.6
Llama3-8B , K=15
0.6
0.5
Llama3-8B , K=20
0.7
0.7
Appendix
Table G.2: Self-consistency threshold τ for SC-BoN-ReAct, tuned on the Dev split.
Test Personas
Method
PSR (%)
#Q
Privileged / Forced
Seen Personas
Always-Ask LLM
52.1 ± 0.9
22.2 ± 0.1
Teacher LLM
65.5 ± 1.1
0.0
Unseen Personas
Always-Ask LLM
50.8 ± 1.6
22.4 ± 0.2
Teacher LLM
65.5 ± 1.9
0.0
Fine-tuned
Appendix
Table I.1: Main results on the ADAPT benchmark (4-fold cross-validation; std across split means). Patel et al. (2025) numbers are reproduced from their Table 1. In-context examples from seen personas injected at hypothesis particle regeneration.
In hierarchical reasoning, failures often originate at intermediate decision points where the agent commits to a wrong branch without recognizing that it lacks critical information. Rather than treating clarification as an external uncertainty trigger, we propose ACTION-RATING, a formulation that places it inside the agent's action space on a shared ordinal scale with navigation, so that asking competes directly with acting at every decision point and help-seeking becomes observable at intermediate states. Two structurally distinct information-seeking modes emerge from the agent's own ratings: mandatory (no viable branch) and opportunistic (residual uncertainty despite a leading candidate). On Harmonized Tariff Schedule classification (30,000-node taxonomy, three benchmarks, 9~LLMs across 4 families), we observe a regime shift from mandatory to opportunistic clarification, with Information-Seeking Effectiveness (ISE), a local diagnostic defined as the fraction of help interactions followed by a correct next navigation step (not a final-task metric), rising from 50% to 74%. Three diagnostic contrasts fail to reproduce this structure. A separability test shows that the information-seeking pattern (mode split, ISE ranking) persists when answer quality is degraded (-18.8% accuracy), supporting an empirical separation between where an agent seeks help and the quality of the help it receives. Under the controlled answer channel, accuracy gains reach +16.2% at 10-digit; we read this as an upper bound on what better localization could unlock, not a deployment estimate.
Humans often specify tasks incompletely, so assistants must know when and how to ask clarifying questions. However, effective clarification remains challenging in software engineering tasks as not all missing information is equally valuable, and questions must target information users can realistically provide. We study clarification in real software engineering tasks by quantifying which types of information most affect task success and which questions elicit useful responses from simulated users. Using Shapley attribution and distributional comparisons, we identify two key properties of effective clarification: task relevance (which information predicts success) and user answerability (what users can realistically provide). We operationalize these properties as multi-stage reinforcement learning rewards to train CLARITI, an 8B-parameter clarification module, that matches GPT-5's resolution rate on underspecified issues while generating 41% fewer questions. Our results suggest that grounding reward design in empirical analysis of information impact and user answerability improves clarification efficiency.
Sanidhya Vijayvargiya, Vijay Viswanathan, Graham Neubig
Language Technologies Institute, Carnegie Mellon University, Pittsburgh, USA.
Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreversible errors. When instructions are incomplete, the agent must decide not only whether to ask for clarification but when, and no prior work measures how clarification value changes over the course of execution. We introduce a forced-injection framework that provides ground-truth clarifications at controlled points in the agent's trajectory across four information dimensions (goal, input, constraint, context), three agent benchmarks, and four frontier models (three per benchmark; one on a single benchmark only; 84 task variants; 6,000+ runs). Counter to the common intuition that "earlier is always better," we find that the value of clarification depends sharply on what information is missing: goal clarification loses nearly all value after 10% of execution (pass@3 drops from 0.78 to baseline), while input clarification retains value through roughly 50%. Deferring any clarification type past mid-trajectory degrades performance below never asking at all. Cross-model Kendall tau correlations (0.78-0.87 among models sharing identical task coverage; 0.34-0.67 across the full 4-model panel) confirm these timing profiles are substantially task-intrinsic. A complementary study of 300 unscripted sessions reveals that no current frontier model asks within the empirically optimal window, with strategies ranging from over-asking (52% of sessions) to never asking at all. These empirical demand curves provide the quantitative foundation that existing theoretical frameworks require but have lacked, and establish concrete design targets for timing-aware clarification policies. Code and data will be publicly released.