Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box2-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box2-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.
Figures & tables
Figure 1 : The top illustration contrasts independent execution with guidance from capable and weak teachers. The bottom panels show performance under self-solving, good-workflow, and bad-workflow conditions on (a) AIME 2026 and (b) WebShop. Good workflows improve performance, while bad ones expose sensitivity to misleading guidance.
Model
Performance
Box 2 -Bench Effects
Base
Good
Bad
Partial
Mixed
Utilization
Robustness
Recovery
S0
SG
SB
SP
SM
Δuse
Δbad
Δstop
Δswitch
# Math — OpenR1-Math
Gemini 3.7 Flash
76.7%
83.3%
33.3%
80.0%
43.3%
+6.7
-43.3
-3.3
-36.7
DeepSeek V4 Flash
76.7%
80.0%
66.7%
83.3%
73.3%
+3.3
-10.0
+3.3
-10.0
GLM 5.2
80.0%
80.0%
40.0%
83.3%
56.7%
0.0
-40.0
+3.3
-26.7
Table 1 : Box 2 performance across five execution regimes. We first report absolute performance under each regime, followed by the corresponding paired workflow effects.
Model
Performance
Box 2 -Bench Effects
Base
Good
Bad
Partial
Mixed
Utilization
Robustness
Recovery
S0
SG
SB
SP
SM
Δuse
Δbad
Δstop
Δswitch
# Math — AIME2026
Qwen3-0.6B
16.7%
23.3%
20.0%
15.6%
15.6%
+6.7
+3.3
-7.8
+0.0
Qwen3-4B
66.7%
80.0%
46.7%
81.1%
66.7%
+13.3
-20.0
+1.1
-14.4
Qwen3-8B
76.7%
83.3%
50.0%
84.4%
74.4%
+6.7
-26.7
+1.1
-10.0
Table 2 : Box 2 -Bench performance across model scales on AIME 2026 and WebShop. Good workflows consistently improve performance, whereas bad workflows generally reduce it. Scaling does not reliably eliminate sensitivity to misleading guidance.
Figure 2 : Workflow recovery across model scales on (a) AIME 2026 and (b) WebShop. For each model, Partial retains the first k good steps, whereas Mix combines k good steps with 8−k bad steps, for k∈{0,2,4,6,8} . Increasing the number of good steps generally narrows the performance gap between the two regimes, although recovery can be non-monotonic.
Figure 3 : Construction of counterfactual supervised fine-tuning data. For each problem, we sample the base model to obtain a verified successful trace, which the planner uses to synthesize paired correct and incorrect workflows. Each training example combines the problem and an incorrect workflow with the verified correct response, providing supervision for completing the task despite misleading procedural guidance.
Model
Performance
Box 2 -Bench Effects
Base
Good
Bad
Partial
Mixed
Utilization
Robustness
Recovery
S0
SG
SB
SP
SM
Δuse
Δbad
Δstop
Δswitch
# Math — AIME 2026
Base
66.7%
80.0%
46.7%
81.1%
66.7%
+13.3
-20.0
+1.1
-14.4
+ SFT
73.3%
70.0%
66.7%
74.4%
70.0%
-3.3
-6.7
+4.4
-4.4
+ SFT + RLenv
73.3%
80.0%
56.7%
73.3%
74.4%
+6.7
-16.7
-6.7
+1.1
Table 3 : Performance before and after workflow training on AIME 2026 and WebShop. SP and SM average the three Partial and Mixed conditions, respectively. Effects report utilization ( SG−S0 ), robustness ( SB−S0 ), and recovery when guidance stops ( SP−SG ) or becomes misleading ( SM−SP ). RLenv uses direct task-outcome rewards on bad-workflow inputs.
Figure 4 : Workflow recovery across training stages on (a) AIME 2026 and (b) WebShop. The columns compare the base model, SFT model, and SFT model followed by RL under the same partial- and mixed-workflow compositions. SFT generally reduces the separation between the two regimes, while the effect of RL depends on the task and workflow composition.
Figure 5 : Selective reliance across sources of context. From left to right, the panels illustrate checking a proposed workflow, revising a peer draft, and verifying a stored claim using tool evidence. The common challenge is to use helpful context without being bound by misleading content. Examples and dialogue are schematic rather than recorded trajectories; correctness markers are explanatory annotations, not model inputs.
Figure 6 : Adaptive EoM evaluation on AIME 2026. In (b), each model’s upper and lower rows show wrong-to-correct and correct-to-wrong revisions, respectively; endpoint colors indicate answer states, and their horizontal separation gives the corresponding rate. The dashed diagonal in (c) marks equal repair and corruption rates. All rates are normalized by 150 episodes (30 problems × 5 episodes).
Model
Base
Good
Bad
Score
Tool Rate
Score
Tool Rate
Score
Tool Rate
# Memory — LongMemEval-V2-Small
Base
21.1%
100.0%
95.6%
72.2%
10.0%
65.6%
+ SFT
28.9%
100.0%
92.2%
72.2%
14.4%
72.2%
+ SFT + RLenv
25.6%
100.0%
94.4%
66.7%
13.3%
74.4%
Table 4 : Memory-conditioned performance on LongMemEval-V2-Small. Score denotes answer accuracy, and Tool Rate denotes the proportion of trajectories invoking the archive tools; Base denotes no memory brief.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Baseline
Full
Partial
Mix
None
Good
Bad
Good-2
Good-4
Good-6
2G+6B
4G+4B
6G+2B
# Search — WebShop
Qwen3.5-4B
20.8%
35.8%
15.2%
28.2%
35.8%
36.0%
27.4%
30.8%
32.4%
Qwen3.5-9B
25.2%
35.6%
15.4%
28.4%
34.8%
35.0%
27.8%
33.0%
37.0%
Qwen3.5-27B
28.6%
40.8%
17.2%
35.6%
40.2%
41.2%
38.2%
39.2%
41.0%
Appendix
Table 5 : WebShop accuracy across all workflow conditions. Good- k retains the first k steps of the good workflow. x G+ y B concatenates the first x good steps with the last y bad steps.
Model
Baseline
Full
Partial
Mix
None
Good
Bad
Good-2
Good-4
Good-6
2G+6B
4G+4B
6G+2B
# Math — AIME2026
Qwen3-0.6B
16.7%
23.3%
20.0%
16.7%
16.7%
13.3%
13.3%
16.7%
16.7%
Qwen3-4B
66.7%
80.0%
46.7%
76.7%
86.7%
80.0%
53.3%
73.3%
73.3%
Qwen3-8B
76.7%
83.3%
50.0%
83.3%
83.3%
86.7%
70.0%
70.0%
83.3%
Appendix
Table 6 : AIME2026 accuracy across all workflow conditions. Good- k retains the first k steps of the good workflow. x G+ y B concatenates the first x good steps with the last y bad steps.
Model
Baseline
Full
Partial
Mix
None
Good
Bad
Good-2
Good-4
Good-6
2G+6B
4G+4B
6G+2B
# Search — WebShop
Base
25.2%
35.6%
15.4%
28.4%
34.8%
35.0%
27.8%
33.0%
37.0%
+ SFT
30.8%
33.6%
30.2%
32.2%
32.4%
33.4%
31.0%
32.2%
32.6%
+ SFT + RLenv
32.0%
34.8%
30.8%
32.8%
34.2%
34.2%
32.4%
32.8%
34.0%
+ SFT + RLrel
29.6%
34.0%
31.4%
32.2%
33.6%
34.0%
32.0%
32.8%
34.0%
Appendix
Table 7 : WebShop success rates across workflow conditions for the outcome-based and relative-reward RL variants. Good- k retains the first k good steps, while x G+ y B uses a length- x good prefix followed by a length- y bad suffix. RLenv uses the WebShop environment return, and RLrel uses a same-snapshot relative reward.
Model
Baseline
Full
Partial
Mix
None
Good
Bad
Good-2
Good-4
Good-6
2G+6B
4G+4B
6G+2B
# Mathematical Reasoning — AIME 2026
Base
66.7%
80.0%
46.7%
76.7%
86.7%
80.0%
53.3%
73.3%
73.3%
+ SFT
73.3%
70.0%
66.7%
73.3%
76.7%
73.3%
63.3%
73.3%
73.3%
+ SFT + RLenv
73.3%
80.0%
56.7%
73.3%
66.7%
80.0%
70.0%
76.7%
76.7%
+ SFT + RLrel
70.0%
80.0%
60.0%
76.7%
73.3%
83.3%
63.3%
66.7%
80.0%
Appendix
Table 8 : AIME 2026 accuracy across workflow conditions for the outcome-based and relative-reward RL variants. Each entry reports majority-vote accuracy over eight samples per problem. Good- k retains the first k good steps, while x G+ y B uses a length- x good prefix followed by a length- y bad suffix. RLenv uses binary answer correctness, and RLrel uses a same-snapshot relative reward.
Figure 7 : Benchmark construction and evaluation pipeline. For each task, we construct good and bad harnesses, verify their structure, review their content with GPT-5.6-sol, and freeze the resulting harnesses before evaluation. We evaluate the model with no harness or a good, partial, mixed, or bad harness to measure utilization ( Δuse ), robustness ( Δbad ), and recovery after a reliability switch ( Δswitch ).
Figure 8 : Single-model workflow case. The task is to count permutations of {1,…,6} whose order divides 6. The shared mixed workflow contains a valid prefix but a misleading suffix that replaces “divides 6” with “equals 6”. The Base model follows the misleading suffix and outputs an incorrect answer, whereas the trained model restores the original condition and recomputes the correct result. Responses shown in the figure are abridged for visualization.
Figure 9 : Multi-agent collaboration case. We compare Base and trained checkpoints on the same AIME 2026 problem under the same A4 → A2 handoff pattern. In both trajectories, the first agent produces an incorrect public draft. The Base second agent continues from the inherited estimate and preserves the wrong answer, whereas the trained second agent re-examines the draft and corrects the answer. Text in the figure is abridged; the thought bubble is a schematic visualization rather than a verbatim model output.
Figure 10 : Memory-augmented reasoning case. Both checkpoints receive the same corrupted memory brief for the same ServiceNow question. The Base model copies the corrupted memory directly, while the trained model verifies the claim against archive evidence and replaces the wrong item with the correct one. The figure shows an abridged view of the interaction; detailed task description is provided in the text.
Language model agents are increasingly effective in solving realistic tasks through multi-turn tool use. However, training reliable tool-using agents remains challenging in practice. While reinforcement learning provides an on-policy paradigm for improving agents from their own environment interactions, its effectiveness depends heavily on the training task distribution. When tasks are fixed before training, the task distribution can become increasingly mismatched with the policy's evolving capabilities, causing many rollouts to be spent on uninformative tasks. We propose SENTINEL, a failure-driven reinforcement learning framework that turns the Solver's rollout failures into targeted training tasks. SENTINEL follows a Controller--Proposer--Solver loop: the Controller analyzes failed trajectories and summarizes recurring error patterns, the Proposer generates executable tasks that stress these weaknesses, and the Solver is trained on the targeted tasks. On Tau2-Bench Retail with Qwen3-4B-Thinking-2507, SENTINEL improves Pass^{}1 from 66.4 to 74.9 and outperforms RL on general synthetic tasks across Pass^{}k metrics. These results demonstrate that model failures provide an effective and scalable source of targeted training signal for improving tool-using language model agents.
Ziyi Wang, Yuxuan Lu, Yimeng Zhang +8
Northeastern University · Independent Researcher · Northwestern University
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misleading signal flips a clean-correct answer to wrong. Across a comprehensive benchmark study, we observe that such a susceptibility is universal. We then propose SCOPE, which mines clean-correct/misleading-wrong failures and optimizes a standard Direct Preference Optimization (DPO) objective over matched preference pairs balanced equally across all four conditions, rather than over misleading items alone. Our approach substantially reduces SC2W on popular open-sourced models while preserving accuracy when the added context is clean, correct, or irrelevant. With this work, we argue that models should be judged on selective trust, not on resistance alone.
Xian Sun, Wei Chow, Yingshuo Wang +4
1Duke University · 2National University of Singapore · 3UC Berkeley +3
Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models produce confident hallucinations, even when wrong answers carry catastrophic consequences. Our evaluations on GSM8K, MedQA and GPQA show frontier models almost never abstain despite explicit warnings of severe penalties, suggesting that prompts cannot override training that rewards any answer over no answer. As a remedy, we propose Reinforced Hesitation (RH): a modification to Reinforcement Learning from Verifiable Rewards (RLVR) to use ternary rewards (+1 correct, 0 abstention, -λ error) instead of binary. Controlled experiments on logic puzzles reveal that varying λ produces distinct models along a Pareto frontier, where each training penalty yields the optimal model for its corresponding risk regime: low penalties produce aggressive answerers, high penalties conservative abstainers. The same frontier holds on MATH Levels 4--5 and on medical QA, where it transfers to an unseen dataset. We then introduce two inference strategies that exploit trained abstention as a coordination signal: cascading routes queries through models with decreasing risk tolerance, while self-cascading re-queries the same model on abstention. Both outperform majority voting with lower computational cost. These results establish abstention as a first-class training objective that transforms ``I don't know'' from failure into a coordination signal, enabling models to earn trust through calibrated honesty about their limits.
Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan Li
Toyota Technological Institute at Chicago · University of California, San Diego