Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents
Organizations: IBM · Weizmann Institute of Science
Abstract
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.
Figures & tables
| quantity | definition |
|---|---|
| required-evidence coverage | share of the required needs recovered, |
| breadth | independent lines in which the run resolves a need at depth |
| depth-weighted recall | with for and 3 to 5 |
| deepest need resolved | , the level the run reached |
| out-of-order rate | share of resolved needs taken before a prerequisite (lower is better) |
| stop-when-done | over the decision points the policy reaches |
| MuSiQue | StrategyQA | 2WikiMultiHopQA | ||||
| Baseline | Q&D | Baseline | Q&D | Baseline | Q&D | |
| Equal spend | ||||||
| Required-evidence coverage (%) | 78.3 | 89.5 | 78.4 | 85.5 | 88.2 | 93.1 |
| Depth-weighted recall (%) | 77.9 | 90.5 | 51.2 | 56.6 | 84.8 | 90.3 |
| Deepest need resolved (depth) | 1.53 | 1.75 | 0.54 | 0.64 | 0.57 | 0.63 |
| Breadth (lines) | 0.87 | 0.99 | 0.47 | 0.52 | 0.77 | 0.84 |
| MuSiQue | StrategyQA | |||||||
| supervision | gain | stop done | ask not done | asks | gain | stop done | ask not done | asks |
| same model, prompted | — | 15.3 | 96.4 | 6.29 | — | 14.6 | 96.3 | 5.59 |
| imitation (first stage) | 95.1 | 86.8 | 3.05 | 90.7 | 83.0 | 1.34 | ||
| question pairs | 98.7 | 86.0 | 2.47 | 98.0 | 83.4 | 1.28 | ||
| stop contrasts alone | 10.1 | 99.6 | 6.95 | 9.7 | 98.1 | 5.75 | ||
| final stage, seed 1 | 42.7 | 94.2 | 4.23 | 54.6 | 91.2 | 2.13 | ||
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| suite | graphs | nodes | edges | depth histogram | behind an edge |
|---|---|---|---|---|---|
| MuSiQue | 800 | 2,660 | 1,860 | 0:1,196 1:800 2:530 3:134 | |
| StrategyQA | 2,290 | 6,720 | 4,483 | 0:3,721 1:2,320 2:615 3:60 4:4 | |
| 2WikiMultiHopQA | 12,576 | 31,120 | 11,956 | 0:19,164 1:11,956 |
| quantity | definition | reads |
|---|---|---|
| required-evidence coverage | evidence recovered | |
| breadth | independent lines touched | |
| depth-weighted recall | , over depths present | coverage per level, weighted deep |
| coverage at depth | one level on its own | |
| deepest need resolved | the level reached at all | |
| out-of-order rate | taken out of order |
| level | name | the needs it pursues | read by |
|---|---|---|---|
| 0 | stated | named in the request | coverage of stated needs |
| 1 | horizontal | at depth zero, unstated | breadth, not separable |
| 2 | vertical, one step | after one prerequisite | depth-weighted recall |
| 3 | vertical, deep | behind two or more prerequisites | deepest need resolved |
| 4 | open | in no annotated graph | not scored by a graph |
| order | taken before its prerequisite | out-of-order rate (lower better) |
| suite | unit | allowance | comparator | comparator | coverage | 95% interval |
|---|---|---|---|---|---|---|
| granted | spend | questions | under budget | |||
| MuSiQue | question tokens | 91.8 | 65.5 | 4.261 | ||
| MuSiQue | question words | 74.3 | 53.3 | 4.250 | ||
| StrategyQA | question tokens | 62.2 | 42.0 | 2.974 | ||
| StrategyQA | question words | 48.2 | 33.2 | 2.881 |
| arm | fitted from | training records | what it varies |
|---|---|---|---|
| imitation reference | base model | decisions | nothing (the comparison point) |
| stop-weighted imitation | base model | decisions | stop rows down-weighted to |
| question pairs | imitation reference | question pairs of a -pair export | preference on question pairs |
| stop contrasts only | imitation reference | pairs, one kind | preference on stop contrasts only |
| both kinds, from reference | imitation reference | pairs, three kinds | preference on both kinds at once |
| both kinds, from pairs | question pairs | pairs, three kinds | the same set, from a trained start |
| supervision | MuSiQue | StrategyQA | 2WikiMultiHopQA |
|---|---|---|---|
| imitation reference | |||
| filtered demonstrations | |||
| stop-weighted imitation | |||
| the same, stop rows down-weighted | |||
| question pairs | |||
| question pairs, from the reference |
| MuSiQue | StrategyQA | 2WikiMultiHopQA | |
| Equal spend: both arms at the lower of their two question counts | |||
| required-evidence coverage | |||
| breadth | |||
| depth-weighted recall | |||
| deepest need resolved | ∘ | ||
| out-of-order rate † | ∘ | ∘ | |
| suite | seed 1 | seed 2 | seeds 1+2 | |
|---|---|---|---|---|
| required-evidence coverage | MuSiQue | |||
| StrategyQA | ||||
| 2WikiMultiHopQA | ||||
| breadth | MuSiQue | |||
| StrategyQA | ||||
| 2WikiMultiHopQA |
| suite | split | 95% interval | tasks | own-stop | |
|---|---|---|---|---|---|
| MuSiQue | development | ||||
| MuSiQue | held out | ||||
| StrategyQA | development | ||||
| StrategyQA | held out | ||||
| 2WikiMultiHopQA | held out |
| criterion | gated | MuSiQue | StrategyQA | 2WikiMultiHopQA | ||
|---|---|---|---|---|---|---|
| dev | held out | dev | held out | held out | ||
| required-evidence coverage | yes | |||||
| question diversity, seed 1 | yes | |||||
| question diversity, seed 2 | yes | |||||
| question length | yes | fail | fail | fail | fail | fail |
| malformed output | yes | |||||
| control | at most the trained questioner’s spend † | shared cap of eight | |
|---|---|---|---|
| answer length | 375 | ||
| compute budget | 375 | ||
| random questions | 376 | ||
| state-blind template | 400 | ||
| depth one ∗ | 400 | ||
| evidence blind ∗ | 400 |
| arm | words | comparator | relative 95% CI | StrategyQA | MuSiQue |
|---|---|---|---|---|---|
| imitation reference | 13.80 | 12.99 | pass | pass | |
| stop contrasts alone | 16.59 | 12.99 | fail | pass | |
| both kinds | 14.31 | 12.99 | fail | pass | |
| both kinds, from the question-pair arm | 22.57 | 12.99 | fail | fail | |
| stop rows down-weighted to | 14.36 | 12.99 | fail | pass |
| composition | pairs | persist. | stop | cov. | 95% CI | div. | asks |
| MuSiQue , 132 tasks | |||||||
| imitation reference | none | 0.868 | 0.951 | 0.830 | 3.05 | ||
| stop contrasts alone | 18,919 | 0.996 | 0.101 | 0.447 | 6.95 | ||
| both kinds | 31,473 | 0.890 | 0.926 | 0.801 | 3.36 | ||
| both kinds, from the question-pair arm | 31,473 | 0.944 | 0.427 | 0.834 | 4.24 | ||
| StrategyQA , 333 tasks | |||||||
| suite | node | implied | observed | detectable | stronger reader |
|---|---|---|---|---|---|
| MuSiQue | |||||
| StrategyQA | |||||
| 2WikiMultiHopQA |
| Task success | Follow-up turns | Hit rate | Recall | ||||
|---|---|---|---|---|---|---|---|
| domain | prompt | Seed 1 | Seed 2 | Pooled | Pooled | Pooled | Pooled |
| retail | base | ||||||
| retail | stop | ||||||
| airline | base | ||||||
| domain | prompt | retrieved a record | exact repeat | near-exact repeat |
|---|---|---|---|---|
| retail | base | |||
| retail | stop | |||
| airline | base | |||
| domain | prompt | Follow-up turns | Task success | Hit rate | Questions |
|---|---|---|---|---|---|
| retail | base | ||||
| retail | stop | ||||
| airline | base | ||||
| airline | stop |
| comparison | raw | Cohen’s [95% clustered] | / clusters |
| the two retained members, both panels | 77.8% | 0.555 [0.481, 0.624] | 2,508 / 201 |
| first retained with replaced | 65.2% | 0.304 [0.244, 0.364] | 3,156 / 211 |
| second retained with replaced | 58.2% | 0.163 [0.053, 0.259] | 2,648 / 202 |
| first retained with replacement | 79.0% | 0.581 [0.522, 0.638] | 2,596 / 202 |
| second retained with replacement | 84.1% | 0.682 [0.629, 0.734] | 2,223 / 197 |
| rule with first retained | 61.5% | 0.230 [0.167, 0.294] | 3,480 / 215 |
| question asked of the rater | items | raters agree | chance-corrected |
|---|---|---|---|
| does candidate A reach an unstated need | |||
| does candidate B reach an unstated need | |||
| which candidate is the better move |
| MuSiQue | StrategyQA | 2WikiMultiHopQA | |
| prerequisite edges | |||
| bare task question | |||
| another task’s question (chance) | |||
| queries without the parent’s answer | |||
| queries with the parent’s answer | |||
| queries with a random entity instead |
| base | training seeds | MuSiQue ( ) | StrategyQA ( ) |
|---|---|---|---|
| B | |||
| B | |||
| B |
| run | all pairs | length matched |
|---|---|---|
| 32B, final checkpoint | ||
| 32B, lowest imitation loss | ||
| 8B, same recipe |
| seed 1 | seed 2 | seed 3 | seed 0, retrained | four runs pooled | |
|---|---|---|---|---|---|
| MuSiQue | |||||
| StrategyQA | |||||
| 2WikiMultiHopQA | |||||