cs.AISep 28, 2026

StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents

Authors: Wenle Liao, Zhao Wang, Jingchao Zhang, Jiajie Jin, Yimeng Xu, Zhicheng Dou

Organizations: Gaoling School of Artificial Intelligence Renmin University of China

Abstract

LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interaction histories, making it difficult for agents to track which analytical artifacts remain valid over increasingly long horizons and changing dependencies. Consequently, stale artifacts may be silently inherited, propagating errors to downstream stages. To address this challenge, we propose StateGuard, an analytical-state validity management framework for long-horizon data agents. StateGuard externalizes evolving analytical progress into a state graph containing constraints, versioned variables, intermediate conclusions, and cross-state relations, treating each state as an executable, verifiable, and traceable object rather than textual memory alone. StateGuard maintains state validity through evidence-grounded verification and hierarchical intervention. To equip StateGuard with these capabilities, we first introduce Manager-Oriented Counterfactual Supervision, which constructs 3K state-centric trajectories through counterfactual runtime synthesis to fine-tune StateGuard for state maintenance, verification, and repair. We then apply Validity-Guided Policy Optimization, using runtime validity evidence to provide fine-grained learning signals for protocol correctness, state grounding, and intervention quality. Experiments on three diverse long-horizon data-analysis benchmarks show that StateGuard consistently improves data-agent performance while reducing dependency-induced downstream error propagation, demonstrating the advantages of explicit analytical-state management for reliable long-horizon data analysis.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

    May 28, 2026Kewei Xu, Xiaoben Lu, Shuofei Qiao +4Long-Horizon AgentsData Science Agents

  2. LemonHarness Technical Report

    Jun 23, 2026Kailong Ren, Fubo Sun, Jiachen Liu +18Long-Horizon AgentsTool Invocation