LLM agents combine reasoning, tool use, and persistent memory to support work across tasks by reusing stored operational records as premises for later actions. However, environmental or requirement changes can invalidate these records, while existing action review, provenance tracking, and clarification mechanisms may leave the underlying persistent state uncorrected. Our audit of coding-agent trajectories identifies candidate failure chains in which invalid records are reused, leading to task failures and unsafe modifications. We propose StateWise, a framework for diagnosing and repairing persistent operational state before action execution. StateWise uses record-level counterfactual replanning to identify decision-critical records, then establishes their current validity through reliability checks, read-only verification of machine-observable facts, and targeted clarification of developer-owned intent. Typed evidence grounding binds evidence to specific records and scopes, enabling persistent corrections with repair lineage. The agent then replans from the repaired state, followed by an independent state-action check before execution. We evaluate StateWise on 150 executable coding-agent cases across diverse runtime environments, workspace configurations, and repository settings, complemented by cross-model evaluations. Under corrupted persistent state, StateWise achieves 93.3% overall correctness, compared with 38.7% for the baseline agent, with no unsafe actions. Component ablations, multi-task experiments, and transfer evaluations further demonstrate effective recovery, persistent corrections, and transferability across repositories and tool interfaces.
Figures & tables
Figure 1 . Persistent-state errors and recovery across memory, action, and outcome. The first two scenarios show how outdated movie times and airport gates cause failures in human memory and agent belief state. Given the same outdated branch-policy record, the baseline proposes a prohibited force push, whereas StateWise verifies and repairs the record before replanning to open a policy-compliant pull request. Four scenarios arranged across three horizontal layers labeled memory, action, and outcome. A person remembers a movie time of 3:00 although it has changed to 2:00, arrives at 2:30, and misses the movie. An agent remembers airport gate B12 although the current gate is B26, goes to B12, and misses the flight. A baseline coding agent uses an outdated record stating that the target branch is unprotected and proposes a force push that conflicts with the current branch policy. StateWise starts with the same erroneous record, performs risk checking, state attribution, verification and repair, and replanning, and changes the proposed operation to opening a pull request.
Figure 2 . Observed failure patterns from coding-agent trajectories. Historical operational records may be reused as premises for later decisions even after conflicting evidence appears. The audit retained 40 candidate chains from 100 sampled sessions after checking cited turns and chronology. Failure-type categories may overlap. A three-stage failure chain shows historical records, agent decisions, and conflicts with current evidence. The figure also summarizes the audit funnel, failure types, and observed outcomes.
Figure 3 . The StateWise pipeline. Attribution uses counterfactual replanning and type/scope binding to identify influential and uncertain records. Resolution checks their current validity using available evidence, read-only verification, or targeted developer clarification. Confirmed defects receive scoped repairs that preserve lineage, while valid records remain semantically unchanged. The planner retrieves the updated context and replans. Execution requires resolved premises and an independent state–action, invariant, and permission check. Persistent corrections remain available to later tasks. A three-stage pipeline comprising Attribution, Resolution, and Repair and Action Control. Attribution retrieves candidate records, applies temporary counterfactual probes, and checks type and scope bindings. Resolution classifies records as valid, defective, or unresolved. Confirmed defects are repaired in the persistent store with lineage retained. The updated store supplies context for replanning and later tasks. A final action check permits execution or returns a safe stop. The example corrects a repository record from main being unprotected to main being protected.
Figure 4 . Composition and annotation quality of the evaluation set. The top row summarizes case categories, coverage of persistent state, action links and conflicts, state-issue types, and execution settings. The bottom row shows defect types and annotator agreement on behavior, supporting evidence, and state values. A visualization summarizing the evaluation setup.
Metric
Definition
Goal
OC ( Overall correctness )
Correct and safe outcomes across all episodes: ∣E∣∑e∈Eye .
↑
AC ( Act completion )
Correct and safe completion of tasks requiring an operation: ∣EA∣∑e∈EAye .
↑
CS ( Correct stop )
Correct and safe stopping for tasks requiring a stop: ∣ES∣∑e∈ESye .
↑
UAR ( Unsafe action rate )
Episodes with at least one safety violation: ∣E∣∑e∈Eve .
↓
AP , AR , AF1 ( Attribution precision, recall, and F1 )
Record-level precision, recall, and their harmonic mean: AP=TP+FPTP , AR=TP+FNTP , AF1=AP+AR2APAR .
↑
RR ( Repair retrieval )
Subsequent contexts retrieving the repaired record: ∣L∣∑ℓ∈Lzℓ .
↑
Table 1 . Evaluation metrics for task outcomes, attribution, cross-task reuse, and cost.
Figure 5 . OC, AC, and UAR under three memory conditions, three common intervention strategies, baselines from prior work, and our method. Memory conditions include correct, corrupted, and removed target memory. Simple interventions include bare approval, approval with an action rationale, and clarification. Baselines include MemLineage ( Ouyang and Hou, 2026 ) , Safety Sentry ( Chen et al., 2026b ) , and Ask-or-Assume ( Edwards and Schuster, 2026 ) . Three bar charts show OC, AC, and UAR. Corrupting target memory lowers correctness and completion and increases unsafe actions. Removing target memory partially improves these outcomes. Additional bars show simple interventions, baseline methods, and StateWise, which achieves high correctness and completion with no observed unsafe action.
Figure 6 . OC and UAR for StateWise and baseline interventions under erroneous target memory. Two bar charts compare OC on the left and UAR on the right across baseline interventions and StateWise.
Figure 7 . OC and AC across four planning models on a separate 20-case evaluation set. Two bar charts compare StateWise's OC on the left and AC on the right across GPT-5.6-sol, DeepSeek v4 Pro, Qwen 3.7 Max, and Claude Fable 5.
Figure 8 . OC and AC of the full StateWise pipeline and six ablation variants on the same 150 cases. Two bar charts compare OC on the left and AC on the right for the full StateWise pipeline and six component-removal variants.
Method
T1
T2–T4
AFS
Correct Memory
93.8
91.7
87.5
Corrupted Memory
0.0
1.0
0.0
No Memory
0.0
0.0
0.0
MemLineage
0.0
1.0
0.0
Safety Sentry
0.0
2.1
0.0
Ask-or-Assume
9.4
6.2
3.1
Table 2 . Correct and safe task outcomes and AFS across four-task sequences (%).
Figure 9 . Recovery across seen and unseen configurations with an unchanged StateWise core and supplied interface contracts. Panels report OC, AC, distinct interface units, contract support, and UAR. Four panels summarize StateWise's OC, AC, interface coverage, contract support, and UAR across seen and unseen configurations.
Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes what a model reads as useless, and bounds the context at little cost. However, it decides from the text of the history alone and sees nothing of how the code is connected. Since a coding agent edits code many times over a single task, and each write can change what code elsewhere means, such maintenance may keep records a write has falsified, drop ones that still hold, and miss code the agent needs next. To overcome these challenges, this paper proposes StateTape, a novel and scalable framework that rewrites a coding agent's context as the repository changes rather than as the context grows. The key idea of StateTape is to model the repository as a symbol-level code graph, whose dependencies and language rules expose which symbols a write can affect. Upon this graph, a tape marks the symbols each write changed, which turns staleness from an inference about text into an observation of the agent's writes. We propose a per-write procedure in which the tape nominates the records a write could have falsified while a small manager model settles what the write log cannot, and further provide a theoretical analysis and TraceBench, a benchmark that labels what an agent is holding against what is actually needed. Empirically, we demonstrate that StateTape can effectively clear falsified records and retrieve what is needed, and thus achieve a higher resolve rate in all experiments spanned by six coding agents and three edit-heavy benchmarks with little computational overhead.
LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or recovering from erroneous ones. Doing so safely requires coordinated actions, yet current agent harnesses lack unified abstractions and mechanisms for managing local and remote state consistently and efficiently. We describe Planarian, an agent runtime with state management that enables agents and users to recover from erroneous actions and explore alternative executions over consistent local and remote environment state. Planarian introduces the abstraction of agent statepoints, which are consistent, restorable point-in-time versions of the environment state. Planarian exposes three state-management primitives to agents and users: (i) snapshot creates a new statepoint spanning local and remote state without requiring external services to support checkpoints: it relies on efficient incremental process and file system snapshotting to capture local sandboxed state, and transparently records compensating actions to undo remote state changes; (ii) rollback restores the environment to a previous statepoint by reverting to a prior local checkpoint and replaying compensating actions for remote state changes; and (iii) fork creates multiple isolated branches from a statepoint, enabling the agent to explore alternatives in parallel. We show that Planarian enables agents to undo mistakes and explore alternatives in parallel, improving task quality by up to 15x, and allows users to recover from erroneous actions with only 3% overhead.
Jinnan Guo, Hao Mark Chen, Kapil Vaswani +2
Imperial College London · Indian Institute of Science Bangalore · Microsoft Security Response Center
Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with known dependencies and varies information demand, retention, and recovery separately from the difficulty of individual operations. Across four models, restoring a missing result raises accuracy on affected recall steps to 100%, compared with 0% for equal-length irrelevant information. Sufficient storage alone does not ensure success: retention policies can discard required results, errors can propagate through later computations, and agents can stop before recovery is complete. We also introduce VESTIGE, which uses agent execution traces to construct semantic graphs and measure information demand for real tasks. Across 72,562 software-agent trajectories, VESTIGE reveals a steeper distance-related decline in solution-relevant rereading for failed runs (RR 0.951 per distance doubling), while adjusted peak demand alone is not associated with failure. Together, these contributions support evaluating whether agents preserve and recover the information their tasks require.