LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing context management methods can overlook how earlier tool results are used in subsequent execution, leaving needed information out of context. We introduce ContextRender, which manages context through a persistent graph of execution dependencies. We develop Tool-Flow Analysis to track how later operations reuse information from earlier tool results, providing a signal called observed reuse. A renderer combines this signal with recency and semantic relevance to select results within a fixed history budget, retaining omitted results for later use. Across AppWorld and 8-objective QA with three execution models, ContextRender outperforms the evaluated context management baselines using a 6K history budget, well below the models' maximum context windows. Within this budget, it achieves task performance close to or above that of passing the full history while reducing mean inference cost by 10.2%-32.2% relative to Full history. Ablations show that observed reuse improves task performance and retention of results reused later.
Figures & tables
Figure 1: Comparison of context management approaches. ContextRender selects history from a persistent graph using recency, semantic relevance, and observed reuse, retaining unselected results for later use.
Figure 2: Execution trace from an AppWorld task in which the agent settles a shared dinner bill among roommates. Information from the contact result produced at Turn 8 is used by later operations at Turns 9, 15, and 16. At Turn 18, the agent needs Nancy’s email address, which has not been copied forward and remains available only from the Turn 8 result.
Figure 3: ContextRender architecture. The runtime adapter records messages, tool calls, and complete tool results in a persistent context graph. Tool-Flow Analysis records observed reuse as execution dependencies. Before each model invocation, the renderer includes the accumulated execution history if it fits within the history budget B ; otherwise, it selects tool results using observed reuse, recency, and semantic relevance. Omitted results remain stored for later invocations.
Figure 4: Overall performance (AppWorld TGC; QA F1) across all tasks versus mean standardized inference cost on the same tasks that trigger context management at B=6 K. Costs are normalized to the mean cost of Full history ( =1 ). “Recency” denotes Recency + Summary.
Model
Benchmark
Full history
Prune
Compact
Recency + Summary
Semantic Retrieval
ACON
Perf.
Cost
Perf.
Cost
Perf.
Cost
Perf.
Cost
Perf.
Cost
Perf.
Cost
GPT-4.1
AppWorld Normal
↓ 2.3%
↓ 32.2%
↑ 14.4%
↓ 3.2%
↑ 12.4%
↓ 0.8%
↑ 7.6%
↓ 1.6%
↑ 7.6%
↓ 18.4%
↑ 3.3%
↓ 14.9%
AppWorld Challenge
↑ 8.5%
↓ 28.8%
↑ 19.7%
↓ 2.7%
↑ 12.1%
↓ 1.8%
↑ 11.6%
↓ 16.7%
↑ 14.4%
↓ 9.7%
↑ 11.1%
↓ 8.1%
8-objective QA
↑ 4.5%
↓ 10.2%
↑ 3.7%
↓ 2.0%
↑ 2.1%
↓ 4.5%
↑ 9.5%
→ 0.0%
↑ 14.4%
→ 0.0%
↑ 10.4%
↓ 5.7%
Muse Glimmer 30B
AppWorld Normal
↑ 5.8%
↓ 12.4%
↑ 14.2%
↑ 1.7%
↑ 7.4%
↑ 0.8%
↑ 8.2%
↓ 3.2%
↑ 10.7%
↓ 11.1%
↑ 9.8%
↓ 2.4%
AppWorld Challenge
↑ 2.9%
↓ 19.9%
↑ 16.8%
↓ 7.2%
↑ 9.2%
↓ 16.8%
↑ 10.1%
↓ 2.3%
↑ 13.6%
↓ 8.5%
↑ 8.2%
↓ 12.2%
Table 1: Relative changes in task performance and mean inference cost of ContextRender versus each baseline and Full history . Green, red, and gray arrows indicate improvements, regressions, and no change, respectively.
Variant
Performance
Later-Used Result Recall
AppWorld Later-Used Result Recall by Result Age
AppWorld (TGC)
QA (F1)
AppWorld
QA
≤3
4–10
11–25
>25
CR–Recency
71.4
48.5
63.8
86.0
95.7
78.7
37.7
18.5
CR–Relevance
73.2
47.5
71.4
64.2
89.5
81.9
54.1
47.1
CR–Recency + Relevance
72.6
50.8
61.6
76.1
84.4
72.5
43.9
22.7
CR–Full (+Reuse)
75.6
53.1
87.2
91.0
99.3
91.5
79.0
68.5
Table 2: Ablation of ContextRender with GPT-4.1 on AppWorld Test-Normal and 8-objective QA at B=6 K. Recall denotes Later-Used Result Recall. Result age is the number of turns between a tool result’s creation and its subsequent reuse. All values are percentages.
Figure 5: Budget sensitivity with GPT-4.1. B denotes the history budget. “Recency” denotes Recency + Summary.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Selector
Wrec
Wrel
Wreuse
CR–Recency
1.0
0.0
0.0
CR–Relevance
0.0
1.0
0.0
CR–Recency + Relevance
1.0
1.0
0.0
CR–Full (+Reuse)
1.0
1.0
1.0
Appendix
Table 3: Scoring weights used in the selector ablations. Zero disables the corresponding term; retained weights remain unchanged.
Result age at reuse
Events
≤3 turns
713
4–10 turns
1,124
11–25 turns
1,099
>25 turns
238
Total
3,174
Appendix
Table 4: AppWorld reuse events by result age.
Result age at reuse
Events
≤3 turns
224
4–10 turns
70
11–25 turns
41
>25 turns
0
Total
335
Appendix
Table 5: 8-objective QA reuse events by result age.
Method
AppWorld Test-Normal
AppWorld Test-Challenge
8-objective QA
TGC (%)
Cost
Norm.
TGC (%)
Cost
Norm.
F1 (%)
Cost
Norm.
Full history
77.4
0.177
1.000
51.1
0.302
1.000
50.8
0.167
1.000
Prune
66.1
0.124
0.701
46.3
0.221
0.732
51.2
0.153
0.916
Compact (LLM)
67.3
0.121
0.684
49.4
0.219
0.725
52.0
0.157
0.940
Recency + Summary
70.2
0.122
0.689
49.6
0.258
0.854
48.5
0.150
0.898
Semantic Retrieval
70.2
0.147
0.831
48.4
0.238
0.788
46.4
0.150
0.898
Appendix
Table 6: Overall task performance and mean inference cost for GPT-4.1. Performance is measured across all tasks; cost is measured on the same tasks that trigger context management when accumulated execution history exceeds B=6 K. Cost is in reference USD; normalized cost uses uncapped Full history ( =1 ) on the same tasks in each benchmark.
Method
AppWorld Test-Normal
AppWorld Test-Challenge
8-objective QA
TGC (%)
Cost
Norm.
TGC (%)
Cost
Norm.
F1 (%)
Cost
Norm.
Full history
81.5
0.137
1.000
58.3
0.161
1.000
50.4
0.269
1.000
Prune
75.6
0.118
0.861
51.3
0.139
0.863
49.3
0.243
0.903
Compact (LLM)
80.4
0.119
0.869
54.9
0.155
0.963
50.4
0.259
0.963
Recency + Summary
79.8
0.124
0.905
54.4
0.132
0.820
48.2
0.245
0.911
Semantic Retrieval
78.0
0.135
0.985
52.8
0.141
0.876
51.1
0.235
0.874
Appendix
Table 7: Overall task performance and mean inference cost for Muse Glimmer 30B. Performance is measured across all tasks; cost is measured on the same tasks that trigger context management when accumulated execution history exceeds B=6 K. Cost is in reference USD; normalized cost uses uncapped Full history ( =1 ) on the same tasks in each benchmark.
Method
AppWorld Test-Normal
AppWorld Test-Challenge
8-objective QA
TGC (%)
Cost
Norm.
TGC (%)
Cost
Norm.
F1 (%)
Cost
Norm.
Full history
82.7
0.1425
1.000
77.0
0.236
1.000
56.5
0.476
1.000
Prune
81.0
0.1185
0.832
57.6
0.180
0.763
55.2
0.424
0.891
Compact (LLM)
82.1
0.126
0.884
70.7
0.222
0.941
58.2
0.444
0.933
Recency + Summary
82.1
0.117
0.821
71.2
0.179
0.758
58.5
0.435
0.914
Semantic Retrieval
78.0
0.120
0.842
67.1
0.162
0.686
55.3
0.419
0.880
Appendix
Table 8: Overall task performance and mean inference cost for DeepSeek V4 Pro. Performance is measured across all tasks; cost is measured on the same tasks that trigger context management when accumulated execution history exceeds B=6 K. Cost is in reference USD; normalized cost uses uncapped Full history ( =1 ) on the same tasks in each benchmark.
Agent runtime
Insertion point
Integration changes
Used for
opencode
experimental.chat.messages.transform plugin hook; the plugin is specified in opencode.json
The runtime’s own compaction and pruning are disabled so that only ContextRender manages historical context.
Integration only
AppWorld ReAct agent
trimmed_messages property replaced by a mixin on the live agent object
One property. Prompting, parsing, retry logic, tool use, runner, configuration, and CLI are unchanged.
AppWorld
ACON UnifiedAgent (smolagents)
MemoryManager subclass overriding get_conversation_history ; the class reference is rebound before agent construction
Memory class only. Runner, agent, prompts, retriever, and grading are unchanged.
8-objective QA
OpenClaw (CLI in Docker)
OpenAI-compatible proxy registered as a custom provider in openclaw.json
No runtime code is modified; context rendering is applied by the proxy before the request is forwarded.
Sanity check
Meta ARE default agent (GAIA2)
The same proxy, selected with the --provider local --endpoint flags
No runtime code is modified.
Integration check
Appendix
Table 9: Agent runtimes connected to ContextRender . For each runtime, we list the adapter insertion point, the runtime component modified by the integration, and its use in this work. Prompting, tool execution, parsing, and grading are not rewritten.
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on τ3-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.
Long-horizon agents depend on context management: systems compress, summarize, and evict old tokens so tasks can continue beyond finite windows. That is safe only when dropped information is no longer needed or has been internalized. Plans are the stress case: they are written early, used for many steps, and first to be evicted. We introduce replay pairing, a diagnostic that runs the same trajectory with and without the plan in history and measures hidden-state cosine distance. On Llama-3.1-70B, plan signal spikes to 0.453 one step after the plan, then falls 4.1x in a single action-observation step; HotpotQA falls 12.4x. This is evidence that standard LLM agents do not carry plans forward as persistent state, and instead depend on the plan remaining in context. A layer-L32 probe detects this decay as a diagnostic, not as proof that it reads plan content itself. Reasoning models add a measurement confound: their <think> traces re-derive plan content, so standard stripping leaves plan evidence in the stripped condition. We name this the reasoning-trace confound and fix it with strict stripping, which removes prior <think> blocks from the stripped run only. It recovers +163% of the step+1 signal in-sample and +153% held out, while not meaningfully changing non-reasoning Llama (+4.8%). On DeepSeek-R1-Distill-Llama-70B, a Llama-trained probe transfers at AUROC 0.748 (p=6e-4), while R1-specific probes reach 1.000, suggesting R1 encodes plan signal in a different hidden-state direction. Finally, a compression stress test shows the practical cost: naive plan eviction cuts ALFWorld success by 34.7pp, while probe-gated re-surfacing does not recover it. The contribution is a measurement and stress-test framework showing that agent-critical information can be context-resident rather than persistent. Context management is load bearing, but plan protection alone is not enough.
Large language model (LLM) agents often struggle in long-context interactions. As the agent accumulates more interaction history, context management approaches such as sliding window and prompt compression may omit earlier structured information that later steps rely on. Recent retrieval-based memory systems surface relevant content but still overlook the causal and logical structure needed for multi-step reasoning. We introduce ContextWeaver, a selective and dependency-structured memory framework that organizes an agent's interaction trace into a graph of reasoning steps and selects the relevant context for future actions. Unlike prior context management approaches, ContextWeaver supports: (1) dependency-based construction and traversal that link each step to the earlier steps it relies on; (2) compact dependency summarization that condenses root-to-step reasoning paths into reusable units; and (3) a lightweight validation layer that incorporates execution feedback. On the SWE-Bench Verified and Lite benchmarks, ContextWeaver improves performance over a sliding-window baseline in pass@1, while reducing reasoning steps and token usage. Our observations suggest that modeling logical dependencies provides a stable and scalable memory mechanism for LLM agents that use tools.