Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions
Organizations: Nanyang Technological University, Singapore · National University of Singapore, Singapore · CFAR, Agency for Science, Technology and Research, Singapore · IHPC, Agency for Science, Technology and Research, Singapore · Department of Statistics and Operations Research, UNC-Chapel Hill, United States · Carnegie Mellon University, United States
Abstract
An agent whose tool keeps returning nothing useful should stop relying on it. In a retrieval environment with controlled source failures, we separate how agents judge results from what they do. We compare stopping at the same step after longer and shorter runs of results the agent judged useless; this contrast is zero for clock- or deadline-driven stopping. Where we record their judgments, the seven agents we test call a failing source's results useless 97-100% of the time, yet most of them rarely stop on that judgment. Prompt cues change when they stop but not what they stop on. Permission to answer from memory and a reasoning mode can bring early stops regardless of evidence, a stated budget moves the 7-8B models' stops to the deadline, and a stopping rule or call cost in the prompt is followed at most partly. Stopping follows the evidence only when the harness enforces an integration step that makes the agent answer after five consecutive results it judged useless. This step raises failing-source success for every model, keeps the stopping point fixed when the budget doubles, and needs no extra judgment call when the agent states its judgments. A pre-registered replication on 300 fresh questions confirms the dissociation and the rule's effect.
Figures & tables
| Condition | What changes | Who stops |
| none | base prompt | agent |
| permit | may answer from memory | agent |
| budget | permit + budget, counter | agent |
| stated rule | permit + rule in words | agent |
| call cost | budget + price per call | agent |
| decide | running count shown | agent |
| Condition | Qwen2.5 7B | Llama-3.1 8B | Qwen3 8B | Qwen3 32B |
| none | .192 | .263 | .448 | .509 |
| permit | .213 | .269 | .423 | .517 |
| budget | .328 | .312 | .464 | .581 |
| stated rule | .206 | .267 | .439 | .526 |
| call cost | .326 | .317 | .458 | .575 |
| decide | .360 | .266 | .426 | – |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Claim (main text) | Pre-registered tests (outcome) | Exploratory | Appendix |
| Agents judge a failing source’s results useless (§ 3 ) | 16b: stated and side-channel judgments agree (pass) | judgment rates; specificity on real results; accuracy of stated judgments | B.2 , C.3 , D.4 |
| Unaided agents rarely answer after five useless judgments (§ 3 ) | J1 (reported); F5, fresh300 (pass 3/3); T2, stated judgments (pass 2/2); R3, Search-R1 (fail) | whether stopping would have paid | C.1 , D.4 , D.2 |
| Without enforcement, stopping does not follow the useless run (§§ 3 – 4 ) | J2: none, permit, budget (pass 12/12), stated rule (fail for Llama-3.1-8B, Qwen3-32B); C, call cost (pass 3/4); PP1, stated rule (partial for Llama-3.1-8B only); H1 and H3 of A20, Claude Haiku 4.5 (partial; not on the evidence); F1–F2, fresh300 (pass 8/8); T1 of A18, stated judgments (pass 4/4) | hazard model; decide | C.1 , C.2 , D.4 , D.3 |
| A stated budget moves stopping to the deadline (§ 4 ) | H9, H10, F-H9, F-H10, D1, D3, D4, Rp3 (pass); L-H9 (fail) and A9 (mixed), Qwen3-32B | D.1 , D.6 | |
| The enforced rule raises success (§ 5 ) | H1, F-H1, L-H1, Rp1, F4, R2, S-H2 (pass); S-H1, Claude Sonnet 5 (fail, ) | D.1 , B.4 , D.4 , D.6 | |
| without loss where the source recovers | H3 (14/16; 18/20 with Qwen3-32B’s four under L-H1), F-H3 (12/12); 13c, recovery after four or five failures (mixed) | D.1 , D.5 |
| Doc | ID | Prediction; models; decision rule | Outcome |
| Orig. | H1 | Rule none (mean6); 7–8B, Claude Haiku 4.5; one-sided, Holm | pass 4/4: / / ; Claude Haiku 4.5 |
| H2 | Rule rule fed a rate-matched random signal (A1); 7–8B; one-sided, Holm | pass 3/3: ( ) / / | |
| H3 | Rule none on rec1, rec2, rec3, clean; 7–8B, Claude Haiku 4.5; 95% LB | 14/16; fail Qwen3-8B rec1 (LB ), rec3 (LB ) | |
| H4 | Decide rule (mean6); 7–8B; one-sided | rule decide / / : fail for Qwen2.5-7B | |
| H5 e | Rule vs. rule fed a lexical detector; 7–8B | / / | |
| H6 e | Thinking vs. rule, incl. rec3; Qwen3-8B | [ , ]; rec3: .437 vs. rule .470 |
| Doc | ID | Prediction; models; decision rule | Outcome |
| A10 | Rp1 | Replicate (new seeds, new draw of failed observations): rule none (mean6); 7–8B; one-sided, Holm | pass 3/3: / / |
| Rp2 | Replicate: rule LB ; none ; 7–8B | rule .19 [.14, .24] / .46 / .46: fail for Qwen2.5-7B; none / / : pass | |
| Rp3 | Replicate: budget arm’s median answering action on persistent and ; 7–8B | pass 3/3: median 8 / 8 / 8; / / | |
| Rp4 | Replicate: combo rule and budget (95% LB); 7–8B | 5/6: vs. rule / / [LB ], fail for Qwen3-8B (repaired: pass); vs. budget LB / / | |
| A11 | PP1 | Stopping rule stated in the prompt: executed on the evidence if LB , persistent median and non-inferior to the rule; not executed if CI includes or lies below 0 or persistent answer rate ; else partial; 7–8B, Qwen3-32B | not executed / partial / not executed; Qwen3-32B not executed (design-label / / ; ) |
| PP2 | Stated rule vs. external rule (non-inferiority, margin ) and vs. permit (mean6) | fail 4/4 vs. rule: / / ; Qwen3-32B (LB / / / ; all upper bounds ; repaired: Qwen3-8B [ , ], others unchanged in sign); vs. permit / / ; |
| Doc | ID | Prediction; models; decision rule | Outcome |
| A16 | F1–F2 | fresh300: own-judgment UB for none and budget; 7–8B (Qwen3-32B if time allows) | pass 8/8: none / / , Qwen3-32B ; budget / / , Qwen3-32B (UB ; others ) |
| F3 | fresh300: rule LB | pass 4/4: [LB ] / [ ] / [ ]; Qwen3-32B [ ] | |
| F4 | fresh300: rule none (mean6), one-sided, Holm | pass 4/4: / / ; Qwen3-32B (all ) | |
| F5 | fresh300: unaided agent answers on of questions after five own useless judgments (persistent) | pass 3/3: 3/125, 2/283, 0/289 | |
| 16b | Stated vs. side-channel judgments: annotator binary on the adjudicated labels to proceed; agreement expected | .874 [.798, .939]; agreement .87 / .94 / .94; Qwen3-32B .96 | |
| A17 | V | Search-R1: judgment validity, USELESS after failed and after real , else only descriptive | fail : 1.00 and .89 |
| Model | none | permit | budget | stated | cost | decide | enforced | combo |
| Qwen2.5-7B | ||||||||
| Llama-3.1-8B | ||||||||
| Qwen3-8B | ||||||||
| Qwen3-32B | – | |||||||
| Claude Haiku 4.5 | – | – | – | – | – | – | ||
| Claude Sonnet 5 † | – | – | – | – | – | – |
| HotpotQA test300 | FEVER | |||||||||||
| raw | repaired | raw | repaired | |||||||||
| Contrast (mean6) | Q2.5 | L | Q3 | Q2.5 | L | Q3 | Q2.5 | L | Q3 | Q2.5 | L | Q3 |
| rule none (H1, F-H1) | ||||||||||||
| rule random (H2, F-H2) | ||||||||||||
| rule decide (H4) | ||||||||||||
| rule permit (H7) | ||||||||||||
| success per regime | integration index | ||||||||||||
| Model | Arm | pers. | rec1 | rec2 | rec3 | late | clean | plaus. | mean6 | 95% CI | |||
| Qwen2.5-7B | none | .043 | .160 | .153 | .147 | .220 | .427 | .077 | .192 | .02 | .34 | ||
| permit | .033 | .170 | .207 | .137 | .303 | .430 | .143 | .213 | .03 | .32 | |||
| budget | .090 | .407 | .377 | .327 | .333 | .433 | .183 | .328 | .02 | .04 | |||
| decide | .183 | .370 | .300 | .277 | .517 | .513 | .230 | .360 | .03 | .36 | |||
| random | .080 | .187 | .183 | .163 | .300 | .443 | .143 | .226 | .23 | .33 | |||
| success per regime | integration index | |||||||||||
| Model | Arm | pers. | rec1 | rec2 | rec3 | late | clean | mean6 | 95% CI | |||
| Qwen2.5-7B | none | .127 | .675 | .664 | .644 | .572 | .716 | .566 | .01 | .13 | ||
| permit | .151 | .579 | .541 | .538 | .729 | .805 | .557 | .02 | .23 | |||
| budget | .384 | .818 | .764 | .798 | .750 | .801 | .719 | .00 | .03 | .03 | ||
| decide | .668 | .822 | .733 | .729 | .870 | .856 | .780 | .17 | .43 | |||
| random | .377 | .723 | .695 | .682 | .699 | .733 | .651 | .12 | .25 | .14 | ||
| HotpotQA | FEVER | |||||||
| Model | Contrast | mean6 | 95% CI | mean6 | 95% CI | |||
| Qwen2.5-7B | rule none | .001 | .001 | |||||
| rule permit | .055 | .001 | ||||||
| rule budget | 1.000 | .938 | ||||||
| rule random | .029 | .001 | ||||||
| rule lexical | .043 | .001 | ||||||
| 8-action budget | 16-action budget | |||||||||||||||
| Model | Arm | mean3 | pers. | ans. | med. | final | calls | mean3 | pers. | ans. | med. | final | calls | |||
| Qwen2.5-7B | none | .208 | .043 | 156 | 2 | .00 | 3.9 | .219 | .043 | 158 | 2 | .01 | 7.1 | |||
| budget | .300 | .090 | 175 | 8 | .75 | 6.8 | .312 | .093 | 232 | 16 | .75 | 13.4 | ||||
| rule | .246 | .100 | 262 | 2 | .00 | 2.8 | .09 | .249 | .093 | 260 | 2 | .00 | 2.8 | .10 | ||
| combo | .357 | .163 | 282 | 6 | .00 | 4.7 | .45 | .357 | .130 | 267 | 6 | .00 | 4.7 | .44 | ||
| Llama-3.1-8B | none | .241 | .000 | 4 | 4.5 | .25 | 7.9 | .00 | .298 | .010 | 16 | 13 | .06 | 15.6 | .00 | |
| Model | budget | rule | combo |
| Qwen2.5-7B | 5.09 8.69 | 2.88 2.91 | 4.06 4.10 |
| Llama-3.1-8B | 5.45 8.81 | 5.11 5.39 | 4.74 4.99 |
| Qwen3-8B | 4.11 6.62 | 4.10 4.27 | 3.54 3.62 |
| Model | Condition | mean6 | tool calls | extra calls | utility ( ) |
| Qwen2.5-7B | none | 0.192 | 3.63 | 0.00 | |
| budget | 0.328 | 4.90 | 0.00 | ||
| enforced rule | 0.237 | 2.94 | 2.86 | ||
| combo | 0.375 | 4.03 | 4.00 | ||
| stated-judgment rule | 0.209 | 3.46 | 0.08 | ||
| budget + stated-judgment rule | 0.343 | 4.42 | 0.24 |
| E1 | E2 | E3 | E4 | ||
| Model | success vs none | success vs enforced rule | utility vs enforced rule | combination utility vs budget | fires (pers.) |
| Qwen2.5-7B | [ ] | [ ] | [ ] | [ ] | 0.05 |
| Llama-3.1-8B | [ ] | [ ] | [ ] | [ ] | 0.91 |
| Qwen3-8B | [ ] | [ ] | [ ] | [ ] | 0.97 |
| Qwen3-32B | [ ] | [ ] | [ ] | [ ] | 0.89 |
| Claude Haiku 4.5 | mean6 | unaided | enforced rule | design-label contrast | answered (pers.) | median action |
| stated rule | 0.561 | [ ] | [ ] | [ ] | 0.565 | 7 |
| call cost | 0.663 | [ ] | [ ] | [ ] | 0.985 | 7 |