Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box2-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box2-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.
Figures & tables
Figure 1 : The top illustration contrasts independent execution with guidance from capable and weak teachers. The bottom panels show performance under self-solving, good-workflow, and bad-workflow conditions on (a) AIME 2026 and (b) WebShop. Good workflows improve performance, while bad ones expose sensitivity to misleading guidance.
Model
Performance
Box 2 -Bench Effects
Base
Good
Bad
Partial
Mixed
Utilization
Robustness
Recovery
S0
SG
SB
SP
SM
Δuse
Δbad
Δstop
Δswitch
# Math — OpenR1-Math
Gemini 3.7 Flash
76.7%
83.3%
33.3%
80.0%
43.3%
+6.7
-43.3
-3.3
-36.7
DeepSeek V4 Flash
76.7%
80.0%
66.7%
83.3%
73.3%
+3.3
-10.0
+3.3
-10.0
GLM 5.2
80.0%
80.0%
40.0%
83.3%
56.7%
0.0
-40.0
+3.3
-26.7
Table 1 : Box 2 performance across five execution regimes. We first report absolute performance under each regime, followed by the corresponding paired workflow effects.
Model
Performance
Box 2 -Bench Effects
Base
Good
Bad
Partial
Mixed
Utilization
Robustness
Recovery
S0
SG
SB
SP
SM
Δuse
Δbad
Δstop
Δswitch
# Math — AIME2026
Qwen3-0.6B
16.7%
23.3%
20.0%
15.6%
15.6%
+6.7
+3.3
-7.8
+0.0
Qwen3-4B
66.7%
80.0%
46.7%
81.1%
66.7%
+13.3
-20.0
+1.1
-14.4
Qwen3-8B
76.7%
83.3%
50.0%
84.4%
74.4%
+6.7
-26.7
+1.1
-10.0
Table 2 : Box 2 -Bench performance across model scales on AIME 2026 and WebShop. Good workflows consistently improve performance, whereas bad workflows generally reduce it. Scaling does not reliably eliminate sensitivity to misleading guidance.
Figure 2 : Workflow recovery across model scales on (a) AIME 2026 and (b) WebShop. For each model, Partial retains the first k good steps, whereas Mix combines k good steps with 8−k bad steps, for k∈{0,2,4,6,8} . Increasing the number of good steps generally narrows the performance gap between the two regimes, although recovery can be non-monotonic.
Figure 3 : Construction of counterfactual supervised fine-tuning data. For each problem, we sample the base model to obtain a verified successful trace, which the planner uses to synthesize paired correct and incorrect workflows. Each training example combines the problem and an incorrect workflow with the verified correct response, providing supervision for completing the task despite misleading procedural guidance.
Model
Performance
Box 2 -Bench Effects
Base
Good
Bad
Partial
Mixed
Utilization
Robustness
Recovery
S0
SG
SB
SP
SM
Δuse
Δbad
Δstop
Δswitch
# Math — AIME 2026
Base
66.7%
80.0%
46.7%
81.1%
66.7%
+13.3
-20.0
+1.1
-14.4
+ SFT
73.3%
70.0%
66.7%
74.4%
70.0%
-3.3
-6.7
+4.4
-4.4
+ SFT + RLenv
73.3%
80.0%
56.7%
73.3%
74.4%
+6.7
-16.7
-6.7
+1.1
Table 3 : Performance before and after workflow training on AIME 2026 and WebShop. SP and SM average the three Partial and Mixed conditions, respectively. Effects report utilization ( SG−S0 ), robustness ( SB−S0 ), and recovery when guidance stops ( SP−SG ) or becomes misleading ( SM−SP ). RLenv uses direct task-outcome rewards on bad-workflow inputs.
Figure 4 : Workflow recovery across training stages on (a) AIME 2026 and (b) WebShop. The columns compare the base model, SFT model, and SFT model followed by RL under the same partial- and mixed-workflow compositions. SFT generally reduces the separation between the two regimes, while the effect of RL depends on the task and workflow composition.
Figure 5 : Selective reliance across sources of context. From left to right, the panels illustrate checking a proposed workflow, revising a peer draft, and verifying a stored claim using tool evidence. The common challenge is to use helpful context without being bound by misleading content. Examples and dialogue are schematic rather than recorded trajectories; correctness markers are explanatory annotations, not model inputs.
Figure 6 : Adaptive EoM evaluation on AIME 2026. In (b), each model’s upper and lower rows show wrong-to-correct and correct-to-wrong revisions, respectively; endpoint colors indicate answer states, and their horizontal separation gives the corresponding rate. The dashed diagonal in (c) marks equal repair and corruption rates. All rates are normalized by 150 episodes (30 problems × 5 episodes).
Model
Base
Good
Bad
Score
Tool Rate
Score
Tool Rate
Score
Tool Rate
# Memory — LongMemEval-V2-Small
Base
21.1%
100.0%
95.6%
72.2%
10.0%
65.6%
+ SFT
28.9%
100.0%
92.2%
72.2%
14.4%
72.2%
+ SFT + RLenv
25.6%
100.0%
94.4%
66.7%
13.3%
74.4%
Table 4 : Memory-conditioned performance on LongMemEval-V2-Small. Score denotes answer accuracy, and Tool Rate denotes the proportion of trajectories invoking the archive tools; Base denotes no memory brief.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Baseline
Full
Partial
Mix
None
Good
Bad
Good-2
Good-4
Good-6
2G+6B
4G+4B
6G+2B
# Search — WebShop
Qwen3.5-4B
20.8%
35.8%
15.2%
28.2%
35.8%
36.0%
27.4%
30.8%
32.4%
Qwen3.5-9B
25.2%
35.6%
15.4%
28.4%
34.8%
35.0%
27.8%
33.0%
37.0%
Qwen3.5-27B
28.6%
40.8%
17.2%
35.6%
40.2%
41.2%
38.2%
39.2%
41.0%
Appendix
Table 5 : WebShop accuracy across all workflow conditions. Good- k retains the first k steps of the good workflow. x G+ y B concatenates the first x good steps with the last y bad steps.
Model
Baseline
Full
Partial
Mix
None
Good
Bad
Good-2
Good-4
Good-6
2G+6B
4G+4B
6G+2B
# Math — AIME2026
Qwen3-0.6B
16.7%
23.3%
20.0%
16.7%
16.7%
13.3%
13.3%
16.7%
16.7%
Qwen3-4B
66.7%
80.0%
46.7%
76.7%
86.7%
80.0%
53.3%
73.3%
73.3%
Qwen3-8B
76.7%
83.3%
50.0%
83.3%
83.3%
86.7%
70.0%
70.0%
83.3%
Appendix
Table 6 : AIME2026 accuracy across all workflow conditions. Good- k retains the first k steps of the good workflow. x G+ y B concatenates the first x good steps with the last y bad steps.
Model
Baseline
Full
Partial
Mix
None
Good
Bad
Good-2
Good-4
Good-6
2G+6B
4G+4B
6G+2B
# Search — WebShop
Base
25.2%
35.6%
15.4%
28.4%
34.8%
35.0%
27.8%
33.0%
37.0%
+ SFT
30.8%
33.6%
30.2%
32.2%
32.4%
33.4%
31.0%
32.2%
32.6%
+ SFT + RLenv
32.0%
34.8%
30.8%
32.8%
34.2%
34.2%
32.4%
32.8%
34.0%
+ SFT + RLrel
29.6%
34.0%
31.4%
32.2%
33.6%
34.0%
32.0%
32.8%
34.0%
Appendix
Table 7 : WebShop success rates across workflow conditions for the outcome-based and relative-reward RL variants. Good- k retains the first k good steps, while x G+ y B uses a length- x good prefix followed by a length- y bad suffix. RLenv uses the WebShop environment return, and RLrel uses a same-snapshot relative reward.
Model
Baseline
Full
Partial
Mix
None
Good
Bad
Good-2
Good-4
Good-6
2G+6B
4G+4B
6G+2B
# Mathematical Reasoning — AIME 2026
Base
66.7%
80.0%
46.7%
76.7%
86.7%
80.0%
53.3%
73.3%
73.3%
+ SFT
73.3%
70.0%
66.7%
73.3%
76.7%
73.3%
63.3%
73.3%
73.3%
+ SFT + RLenv
73.3%
80.0%
56.7%
73.3%
66.7%
80.0%
70.0%
76.7%
76.7%
+ SFT + RLrel
70.0%
80.0%
60.0%
76.7%
73.3%
83.3%
63.3%
66.7%
80.0%
Appendix
Table 8 : AIME 2026 accuracy across workflow conditions for the outcome-based and relative-reward RL variants. Each entry reports majority-vote accuracy over eight samples per problem. Good- k retains the first k good steps, while x G+ y B uses a length- x good prefix followed by a length- y bad suffix. RLenv uses binary answer correctness, and RLrel uses a same-snapshot relative reward.
Figure 7 : Benchmark construction and evaluation pipeline. For each task, we construct good and bad harnesses, verify their structure, review their content with GPT-5.6-sol, and freeze the resulting harnesses before evaluation. We evaluate the model with no harness or a good, partial, mixed, or bad harness to measure utilization ( Δuse ), robustness ( Δbad ), and recovery after a reliability switch ( Δswitch ).
Figure 8 : Single-model workflow case. The task is to count permutations of {1,…,6} whose order divides 6. The shared mixed workflow contains a valid prefix but a misleading suffix that replaces “divides 6” with “equals 6”. The Base model follows the misleading suffix and outputs an incorrect answer, whereas the trained model restores the original condition and recomputes the correct result. Responses shown in the figure are abridged for visualization.
Figure 9 : Multi-agent collaboration case. We compare Base and trained checkpoints on the same AIME 2026 problem under the same A4 → A2 handoff pattern. In both trajectories, the first agent produces an incorrect public draft. The Base second agent continues from the inherited estimate and preserves the wrong answer, whereas the trained second agent re-examines the draft and corrects the answer. Text in the figure is abridged; the thought bubble is a schematic visualization rather than a verbatim model output.
Figure 10 : Memory-augmented reasoning case. Both checkpoints receive the same corrupted memory brief for the same ServiceNow question. The Base model copies the corrupted memory directly, while the trained model verifies the claim against archive evidence and replaces the wrong item with the correct one. The figure shows an abridged view of the interaction; detailed task description is provided in the text.