Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed. We frame recovery as a causal decision problem. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile. On long-horizon ALFWorld tasks with Qwen3-14B, CIR raises success from 70.33% to 73.33%, a gain of 3.00 percentage points. It leaves all evaluated trajectories with correct observations untouched. Additional controls show that the benefit of recovery cannot be explained solely by the new observation returned by the environment. These results provide a practical way to evaluate recovery and apply it selectively.
Figures & tables
Figure 1: Overview of our framework. (A) We introduce controlled observation errors and try recovery at different times. (B) From the same execution point, paired runs with and without refresh reveal rescue, harm, and unchanged outcomes. (C) CIR uses information available before intervention to decide whether the expected benefit of refresh outweighs its risk.
Figure 2: Paired recovery effects in the 93-task factual-success cohort across delays (a) and in the complete prefix-feasible primary ( n=106 ) and independent ( n=100 ) cohorts (b). “Stale” denotes the two-step stale condition. Error bars are task-level cluster-bootstrap 95% confidence intervals.
Figure 3: Mechanism controls under clean observations at delay 0 (a) and two-step stale observations at delay 4 (b). Points are effects relative to no recovery. Error bars are task-level cluster-bootstrap 95% confidence intervals.
Figure 4: Selective recovery on 75 held-out prefix-feasible tasks. Each task is evaluated under four observation conditions, giving 300 episodes. Points show effects relative to never refreshing. Error bars are task-level cluster-bootstrap 95% confidence intervals. CIR intervenes in 16.3% of episodes.
Observation condition
Never
CIR
Effect (pp, 95% CI)
Rescue / harm
Clean
81.33%
81.33%
0.00[0.00,0.00]
0 / 0
One-step stale
68.00%
69.33%
+1.33[0.00,4.00]
1 / 0
Two-step stale
60.00%
69.33%
+9.33[1.33,17.33]
9 / 2
Missing
72.00%
73.33%
+1.33[0.00,4.00]
1 / 0
Table 1: Test performance of CIR by observation condition on all 75 held-out prefix-feasible tasks.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Joint recovery outcomes under two-step stale observations at delay d=4 . Rows denote full refresh and columns denote content-ablated refresh, each classified relative to the same no-recovery outcome. The analysis includes 65 tasks with matching execution prefixes and excludes two tasks that terminate successfully before recovery. Eight rescues are shared, two occur only under full refresh, and one occurs only under content-ablated refresh.
Paired comparison
Effect (pp)
95% CI
Full refresh − no recovery
+13.85
[4.62,23.08]
Content-ablated refresh − no recovery
+10.77
[1.54,20.00]
Full refresh − content-ablated refresh
+3.08
[−4.62,10.77]
Appendix
Table 2: Paired success effects on the 65 tasks where both recovery operations can be executed. Intervals use 10,000 task-level cluster-bootstrap resamples. All values are in percentage points.
Figure 6: Paired outcomes at recovery delay d=1 in the 93-task factual-success cohort. Each bar partitions trajectories into stable success, rescue, harm, and stable failure.
Large language model (LLM) agents frequently fail on multi-step tasks involving reasoning, tool use, and environment interaction. While such failures are typically logged or retried heuristically, they contain structured signals about where execution broke down. We introduce CausalFlow, an interventional framework that converts failed agent traces into minimal counterfactual repairs and reusable supervision. CausalFlow models execution traces as sequential chains of dependent steps and computes Causal Responsibility Scores(CRS) via step-level counterfactual intervention to identify failure-inducing steps. For these steps, we generate minimally edited repairs that flip the final outcome to success, producing validated contrastive pairs of the form (wrong step, corrected step). CausalFlow supports two complementary uses: targeted test-time repair that recovers from failures with minimal behavioral drift, and training-time supervision suitable for offline preference optimization or reward modeling. Across four benchmarks spanning mathematical reasoning, code generation, question answering, and medical browsing, CausalFlow converts failed executions into validated minimal repairs with high minimality and causal-consensus scores, and demonstrates that causal attribution is necessary for reliable improvement across diverse agent tasks, outperforming heuristic refinement in complex retrieval settings while producing more localized repairs throughout. These results demonstrate that interventional analysis over structured execution traces provides a principled and scalable mechanism for transforming agent failures into reliability gains and learning-ready supervision.
When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observability) or whether it passed (evaluation), but not which step caused the failure. The obvious heuristics are wrong: the step that executes the harmful action is usually not the step that decided on it, and LLM-judge attribution is correlational and unreliable (state-of-the-art step-level accuracy on the Who&When benchmark is about 14%). We present Causal Agent Replay (CAR), which answers the question by intervention: it models an agent run as a structural causal model, applies a do-operation to a step, and re-executes the trajectory forward under the same stochastic policy, measuring the shift in the outcome distribution. We define an intervention algebra over agent steps, a single-step contrastive estimator whose point-of-commitment rule resolves a confound specific to stochastic run-forward, and a budget-bounded Monte-Carlo Shapley estimator that splits credit across interacting steps. Every effect is reported with confidence intervals. We validate against synthetic structural causal models with planted ground truth: the contrastive estimator recovers the pivotal step, and Shapley recovers a two-step interaction (0.44, 0.45, ~0; efficiency sum 0.909 versus the analytic 0.91). CAR is open source and runs on hosted or free local models.
Large language model (LLM) agents increasingly tackle long-horizon tasks through multi-step environment interaction, yet a single erroneous action can alter subsequent states and observations, causing errors to compound over time. Existing methods either correct the context without repairing altered environment states or restore earlier states while discarding useful experience, making it difficult to both eliminate failure conditions and avoid repeating past mistakes. We argue that reliable recovery should instead be treated as a rollback-boundary control problem that jointly determines when to intervene, where to resume, and what information should survive recovery. Based on this view, we propose Rollback-Induced Reflection (RIR), a unified recovery framework that restores execution to a selected prior state while carrying forward reusable knowledge distilled from the abandoned trajectory to guide subsequent decisions. We further characterize recovery through a unified operator over rollback depth and retained memory, providing a general view of state restoration and knowledge retention. Experiments on three long-horizon benchmarks show that RIR consistently improves average task performance across multiple LLM backbones, with structured reflection memory preserving useful experience and selective rollback enabling efficient recovery.