LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interaction histories, making it difficult for agents to track which analytical artifacts remain valid over increasingly long horizons and changing dependencies. Consequently, stale artifacts may be silently inherited, propagating errors to downstream stages. To address this challenge, we propose StateGuard, an analytical-state validity management framework for long-horizon data agents. StateGuard externalizes evolving analytical progress into a state graph containing constraints, versioned variables, intermediate conclusions, and cross-state relations, treating each state as an executable, verifiable, and traceable object rather than textual memory alone. StateGuard maintains state validity through evidence-grounded verification and hierarchical intervention. To equip StateGuard with these capabilities, we first introduce Manager-Oriented Counterfactual Supervision, which constructs 3K state-centric trajectories through counterfactual runtime synthesis to fine-tune StateGuard for state maintenance, verification, and repair. We then apply Validity-Guided Policy Optimization, using runtime validity evidence to provide fine-grained learning signals for protocol correctness, state grounding, and intervention quality. Experiments on three diverse long-horizon data-analysis benchmarks show that StateGuard consistently improves data-agent performance while reducing dependency-induced downstream error propagation, demonstrating the advantages of explicit analytical-state management for reliable long-horizon data analysis.
Figures & tables
Figure 1: Comparison between our proposed StateGuard and linear reasoning trajectory.
Figure 2: Overview of the StateGuard framework. StateGuard externalizes analytical states during agent execution, performs evidence-grounded validity verification and hierarchical repair, learns from manager-oriented trajectories constructed via relation-based and counterfactual synthesis.
Method
LongDS
DAComp-DE
DABstep
Edu.
Comm.
Soc.
Bus.
Geo.
Spo.
Avg.
DCR
Impl.
Evol.
DCR
Hard
Easy
CFS
CS
CFS
Proprietary General Agents
GPT-5.5
63.56
60.30
33.76
38.60
45.40
41.50
47.14
55.34
24.68
62.14
16.10
53.10
36.77
59.72
+ StateGuard
67.74
65.38
40.57
42.79
56.84
45.62
54.12
51.25
28.91
67.32
21.05
49.96
43.65
61.11
Claude Sonnet 5
64.52
57.48
37.13
32.54
40.29
38.76
44.33
60.56
27.64
55.37
19.41
56.20
32.28
77.78
Table 1: Main results on three data-analysis benchmarks. LongDS reports six-domain performance and average accuracy; DAComp-DE covers data implementation and evolution; DABstep reports hard and easy accuracy. DCR measures excess downstream error associated with incorrect upstream. Bold indicates the best result within each group, and — denotes undefined results.
Figure 3: Left: ablations on training strategy and analytical-state management. Right: Top: task performance v.s. analytical-state coverage. Bottom: verification engagement across stage difficulty.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Hyperparameter
Value
SFT
Backbone
Qwen3-8B
Training trajectories
3K
Epochs
3
Learning rate
5×10−5
Batch size
32
RL
Algorithm
DAPO
Appendix
Table 2: Training configuration and runtime reward settings for StateGuard.
Benchmark
Baseline Worker Budget
+StateGuard Worker Budget
Manager Review
LongDS-Bench
40 / turn
40 / turn
Turn boundary
DABstep
10
15
Every 3 worker steps
DAComp-DE
50
60
Every 5 worker steps
Appendix
Table 3: Worker interaction budgets used for baseline and StateGuard evaluation. Manager actions do not consume worker steps.
Variant
Lifecycle
State
Tools
Repair
Memory
State/Check Policy
ReAct
–
–
–
–
–
–
w/o State Validity
×
×
×
×
×
×
Free-Form Memory
✓
✓
✓
✓
✓
Free
StateGuard
✓
✓
✓
✓
✓
Structured
Appendix
Table 4: Controlled analytical-state ablations. “Free” denotes an available capability without StateGuard’s prescribed state or verification policy.
Benchmark
Reviews / Task
Manager Calls / Task
Time / Manager Call
LongDS-Bench
33.1
134.5
18.4 s
DABstep
2.3
7.3
15.0 s
DAComp-DE (Impl.)
12.0
30.6
31.5 s
DAComp-DE (Evol.)
12.3
35.1
26.9 s
Appendix
Table 5: Runtime characteristics of StateGuard with DeepSeek-V4-Pro workers. Manager inference is served locally with Qwen3-8B.
Figure 4: Internal runtime statistics of StateGuard with DeepSeek-V4-Pro as the worker. Left: state construction statistics, including average states per task, task coverage, and per-state content density. Right: intervention behavior, including the frequency of state hints, error hints, and repairs, as well as tool-usage coverage. Across benchmarks, StateGuard maintains dense structured analytical states while keeping explicit repair interventions sparse and selective.
Figure 5: Representative LongDS case from Global Data on Sustainable Energy . An ambiguity at Turn 3 leads to two downstream paths: the upper path propagates an inconsistent analytical base, whereas StateGuard retrieves the relevant upstream state, prompts localized recomputation, and preserves a consistent state for subsequent ranking.
Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We introduce LongDS, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states. LongDS comprises 68 tasks constructed from real-world Kaggle notebooks, spanning 2,225 turns across six domains including Geoscience, Business, and Education. Tasks are designed around state-evolution patterns (e.g., counterfactual perturbation, rollback, multi-state composition), with an average dependency span of 11.3 turns. Evaluating five state-of-the-art models, we find that the best model reaches only 48.45% average accuracy, performance drops nearly 47 points from early to late turns, and long-horizon errors account for 52%--69% of failures. Further analysis shows that additional agent steps do not necessarily improve performance, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget. We release LongDS to support research on reliable long-horizon agentic data analysis. Code and data will be released at https://github.com/zjunlp/DataMind.
Kewei Xu, Xiaoben Lu, Shuofei Qiao +4
Zhejiang University · Zhejiang University - Ant Group Joint Laboratory of Knowledge Graph · Ant Group
Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench2.1, and from 2.8% to 8.3% on OSWorld2.0. It also raises Claude Opus4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.
As large language model (LLM) agents are applied to longer tasks, they increasingly modify workspace state across multiple rounds of iteration. However, agents typically observe only tool outputs and log fragments, while the actual state changes occur in the file system. Without explicit workspace boundaries, state-changing operations such as file writes and temporary artifact generation may scatter changes across paths. Over time, these weakly constrained changes accumulate, making states such as modified files difficult to track. This paper presents LemonHarness, an integrated execution framework for long-horizon agents. LemonHarness establishes an explicit execution boundary by constraining state-changing operations within a clearly defined workspace and bringing model invocation, tool execution, and rule knowledge within a single controlled boundary. State-changing operations, including file writes, dependency installation, and temporary artifact creation, are executed through structured tool interfaces, with execution feedback recorded as observations available to subsequent model decisions. The system also introduces a reusable rule knowledge base, which turns recurring execution rules and acceptance criteria into runtime knowledge. LemonHarness further adds a time-aware execution mechanism that exposes elapsed and remaining budget to the model, so it can rebalance exploration, implementation, and validation effort as time pressure shifts and avoid timeouts from long waits or excessive verification. On Terminal-Bench 2.0, LemonHarness_GPT-5.3-CodeX reached 84.49% accuracy over 445 trials; pairing the same framework with the stronger GPT-5.5 backbone raised the average accuracy to 86.52% across five jobs. The results suggest that a unified runtime boundary, callable rule knowledge, and time-aware execution can improve the stability of long-horizon agent execution.