When a user changes their mind partway through a task, an agent that has already split the task into sub-goals and paid for tool calls must decide, per cached sub-result, whether to keep, patch, or discard it (salvage), restarting wastes valid work and continuing unchanged answers the old question. Our main finding is that salvage quality is a matter of role design rather than model capability: a language model asked the keep/patch/discard question one node at a time is unreliable, but asked to classify the revision once, with a deterministic layer propagating the decision, it reaches the cost-optimal oracle on all three models tested, from two vendors. Modeling the plan as a commitment hierarchy and the intent change as a belief-revision operator with AGM style postulates, we prove that no policy observing only a node's local view can be both safe and cost optimal, while the single classification design is both. Across three environments the policy recovers the full achievable savings, 43% cheaper than restart, at 100% correctness.
Figures & tables
strategy
labels
corr.
savings
over-salv.
restart
—
1.00
0.00
0
naive-continue
—
0.00
—
—
keep-all
—
0.67
—
360
k/p/d
deterministic rule (= oracle)
1.00
1.00
0
k/p/d
per-node judge
1.00
0.44
0
k/p/d
structured, paraphrased
1.00
0.90 [0.86, 0.95]
0
Table 1: CostBench ( 1,080 instances); k/p/d is keep/patch/discard; savings on the salvageable subset. Judge rows use gpt-4o ; the paraphrased row also hides depth integers ( n=300 , 95% CI). Over-salvage counts stale nodes reused.