Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on τ3-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.
Figures & tables
Bundled Web Shopping
Group Travel Planning
Formal Reasoning
Math
Physics
Method
SR ↑
PS ↑
#Tok. ↓
SR ↑
PS ↑
sPS ↑
#Tok. ↓
SR ↑
PS ↑
#Tok. ↓
SR ↑
PS ↑
#Tok. ↓
Full Context
DeepSeek-V4-Flash
0.00
21.11
52.65M
0.37
3.09
88.43
231.59M
30.00
50.36
40.60M
55.00
72.50
3.68M
GPT-5.6-terra
0.00
33.89
313.95M
–
–
–
–
42.50
54.23
14.59M
50.00
62.60
2.76M
Qwen3.5-397B-A17B
0.00
19.56
134.82M
0.37
6.96
60.99
223.70M
30.00
35.27
19.97M
50.00
55.56
2.02M
Table 1: Results on MemoryArena. Performance metrics are reported in percent; #Tok. denotes total token usage. Bold and underlined performance scores indicate the best and second-best results, respectively. Δ denotes changes relative to DeepSeek-V4-Flash with full context (FC): absolute percentage-point changes for performance and relative percentage changes for token usage (negative values indicate savings).
Airline
Retail
Method
PR ↑
DB ↑
Act. ↑
#Tok. ↓
PR ↑
DB ↑
Act. ↑
#Tok. ↓
Full Context
58.0
70.5
81.4
21.03M
73.7
80.7
90.4
23.55M
FlowState (Ours)
78.0
78.0
83.8
8.16M
81.6
86.0
85.5
18.32M
Δ vs. Full Context
+20.0
+7.5
+2.4
-61.21%
+7.9
+5.3
−4.9
-22.24%
Table 2: Results on τ3 -Bench using DeepSeek-V4-Flash for both methods. PR, DB, and Act. denote pass rate, database accuracy, and action accuracy (%), respectively; #Tok. denotes total token usage. Bold marks better performance. Δ is relative to Full Context: percentage points for performance and percentage changes for tokens (negative means savings).
Travel Planning
Web Shopping
Variant
SR ↑
PS ↑
sPS ↑
#Tok. ↓
SR ↑
PS ↑
#Tok. ↓
FlowState
0.74
12.57
91.87
128.93M
5.33
42.89
44.83M
w/o PSA
0.37
9.36
87.18
160.27M
3.33
44.67
115.37M
w/o ISU
0.00
9.15
90.34
93.16M
0.00
32.11
69.01M
Table 3: Ablation results on Group Travel Planning and Bundled Web Shopping. Performance metrics are in percent; #Tok. denotes total token usage. Bold scores mark the best performance within each environment.
Web Shopping
Travel Planning
Formal Reasoning
Method
nAUC ↑
DR ↓
nAUC ↑
DR ↓
nAUC ↑
DR ↓
Full Context
57.22
10.48
18.35
7.30
69.82
4.63
ReasoningBank
64.71
5.06
7.37
8.43
58.84
7.08
Mem0
47.11
6.09
7.14
8.33
52.41
5.50
BM25
47.97
8.26
10.14
8.94
64.33
6.52
ZipAct
57.17
9.57
8.48
8.54
68.00
5.57
Table 4: Relative performance retention across subtask depths. Metrics use accuracy normalized to each method’s initial value. nAUC measures normalized curve area; DR is the fitted decline rate (percentage points per subtask). Best results are bolded and second-best results underlined.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Web Shopping
Math
Physics
Method
SR ↑
PS ↑
Tok. (M) ↓
SR ↑
PS ↑
Tok. (M) ↓
SR ↑
PS ↑
Tok. (M) ↓
Qwen3.5-397B-A17B
Full Context
0.00
19.56
134.82
30.00
35.27
19.97
50.00
55.56
2.02
FlowState
3.33
37.22
57.68
37.50
53.96
14.53
60.00
59.72
2.01
Δ (pp / %)
+3.33
+17.66
-57.22%
+7.50
+18.69
-27.24%
+10.00
+4.16
-0.50%
DeepSeek-V4-Flash
Appendix
Table 5: FlowState improves task performance and reduces token usage across three base models.
Perspective
State Type
Semantic Scope
User
preferences
User preferences, constraints, and requirements extracted from requests.
Environment
knowledge
Task-relevant observations recorded from tool outputs.
Environment
attributes
Entity-centric information distilled from tool outputs, including properties and characteristics.
Agent
judgements
Agent assessments of task goals, plans, intermediate decisions, and execution risks.
Agent
artifacts
Reusable artifacts, tools, code, and other outputs created by the agent during task execution.