Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on τ3-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.
Figures & tables
Bundled Web Shopping
Group Travel Planning
Formal Reasoning
Math
Physics
Method
SR ↑
PS ↑
#Tok. ↓
SR ↑
PS ↑
sPS ↑
#Tok. ↓
SR ↑
PS ↑
#Tok. ↓
SR ↑
PS ↑
#Tok. ↓
Full Context
DeepSeek-V4-Flash
0.00
21.11
52.65M
0.37
3.09
88.43
231.59M
30.00
50.36
40.60M
55.00
72.50
3.68M
GPT-5.6-terra
0.00
33.89
313.95M
–
–
–
–
42.50
54.23
14.59M
50.00
62.60
2.76M
Qwen3.5-397B-A17B
0.00
19.56
134.82M
0.37
6.96
60.99
223.70M
30.00
35.27
19.97M
50.00
55.56
2.02M
Table 1: Results on MemoryArena. Performance metrics are reported in percent; #Tok. denotes total token usage. Bold and underlined performance scores indicate the best and second-best results, respectively. Δ denotes changes relative to DeepSeek-V4-Flash with full context (FC): absolute percentage-point changes for performance and relative percentage changes for token usage (negative values indicate savings).
Airline
Retail
Method
PR ↑
DB ↑
Act. ↑
#Tok. ↓
PR ↑
DB ↑
Act. ↑
#Tok. ↓
Full Context
58.0
70.5
81.4
21.03M
73.7
80.7
90.4
23.55M
FlowState (Ours)
78.0
78.0
83.8
8.16M
81.6
86.0
85.5
18.32M
Δ vs. Full Context
+20.0
+7.5
+2.4
-61.21%
+7.9
+5.3
−4.9
-22.24%
Table 2: Results on τ3 -Bench using DeepSeek-V4-Flash for both methods. PR, DB, and Act. denote pass rate, database accuracy, and action accuracy (%), respectively; #Tok. denotes total token usage. Bold marks better performance. Δ is relative to Full Context: percentage points for performance and percentage changes for tokens (negative means savings).
Travel Planning
Web Shopping
Variant
SR ↑
PS ↑
sPS ↑
#Tok. ↓
SR ↑
PS ↑
#Tok. ↓
FlowState
0.74
12.57
91.87
128.93M
5.33
42.89
44.83M
w/o PSA
0.37
9.36
87.18
160.27M
3.33
44.67
115.37M
w/o ISU
0.00
9.15
90.34
93.16M
0.00
32.11
69.01M
Table 3: Ablation results on Group Travel Planning and Bundled Web Shopping. Performance metrics are in percent; #Tok. denotes total token usage. Bold scores mark the best performance within each environment.
Web Shopping
Travel Planning
Formal Reasoning
Method
nAUC ↑
DR ↓
nAUC ↑
DR ↓
nAUC ↑
DR ↓
Full Context
57.22
10.48
18.35
7.30
69.82
4.63
ReasoningBank
64.71
5.06
7.37
8.43
58.84
7.08
Mem0
47.11
6.09
7.14
8.33
52.41
5.50
BM25
47.97
8.26
10.14
8.94
64.33
6.52
ZipAct
57.17
9.57
8.48
8.54
68.00
5.57
Table 4: Relative performance retention across subtask depths. Metrics use accuracy normalized to each method’s initial value. nAUC measures normalized curve area; DR is the fitted decline rate (percentage points per subtask). Best results are bolded and second-best results underlined.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Web Shopping
Math
Physics
Method
SR ↑
PS ↑
Tok. (M) ↓
SR ↑
PS ↑
Tok. (M) ↓
SR ↑
PS ↑
Tok. (M) ↓
Qwen3.5-397B-A17B
Full Context
0.00
19.56
134.82
30.00
35.27
19.97
50.00
55.56
2.02
FlowState
3.33
37.22
57.68
37.50
53.96
14.53
60.00
59.72
2.01
Δ (pp / %)
+3.33
+17.66
-57.22%
+7.50
+18.69
-27.24%
+10.00
+4.16
-0.50%
DeepSeek-V4-Flash
Appendix
Table 5: FlowState improves task performance and reduces token usage across three base models.
Perspective
State Type
Semantic Scope
User
preferences
User preferences, constraints, and requirements extracted from requests.
Environment
knowledge
Task-relevant observations recorded from tool outputs.
Environment
attributes
Entity-centric information distilled from tool outputs, including properties and characteristics.
Agent
judgements
Agent assessments of task goals, plans, intermediate decisions, and execution risks.
Agent
artifacts
Reusable artifacts, tools, code, and other outputs created by the agent during task execution.
LLM-based agents increasingly tackle long-horizon tasks with interdependent decisions, where each action reshapes future constraints and intermediate errors can cascade. Existing RAG and agent memory systems organize histories by semantic similarity, retrieving content-relevant entries at decision time. We argue that this design mismatches execution-state dependencies: it fragments decision trajectories and mixes valid and erroneous traces, hindering coherent state reconstruction and error isolation. We propose MAGE (Memory as Agent-Guided Exploration), an active execution-state manager that stores interactions in a hierarchical state tree. The agent derives its state from the active root-to-current path, combining subgoal summaries, recent traces, and hints from prior branches. Four coupled operations maintain the tree: Grow records new traces, Compress summarizes completed subgoals, Maintain validates summaries, and Revise restores a target boundary and resumes on a new branch. This design bounds context growth while preserving state integrity and isolating flawed segments from the active path. Experiments on MemoryArena show that MAGE improves the average task success rate by 7.8--20.4 pp over baselines, while reducing token consumption by 55.1%.
Yaoqi Chen, Haibin Lai, Yuru Feng +12
University of Science and Technology of China · Microsoft · University of California, San Diego +1
Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench2.1, and from 2.8% to 8.3% on OSWorld2.0. It also raises Claude Opus4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.
Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
Yu Luo, Jiamin Jiang, Yimin Zuo +9
Nankai University · Alibaba Group · Tsinghua University