LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing context management methods can overlook how earlier tool results are used in subsequent execution, leaving needed information out of context. We introduce ContextRender, which manages context through a persistent graph of execution dependencies. We develop Tool-Flow Analysis to track how later operations reuse information from earlier tool results, providing a signal called observed reuse. A renderer combines this signal with recency and semantic relevance to select results within a fixed history budget, retaining omitted results for later use. Across AppWorld and 8-objective QA with three execution models, ContextRender outperforms the evaluated context management baselines using a 6K history budget, well below the models' maximum context windows. Within this budget, it achieves task performance close to or above that of passing the full history while reducing mean inference cost by 10.2%-32.2% relative to Full history. Ablations show that observed reuse improves task performance and retention of results reused later.
Figures & tables
Figure 1: Comparison of context management approaches. ContextRender selects history from a persistent graph using recency, semantic relevance, and observed reuse, retaining unselected results for later use.
Figure 2: Execution trace from an AppWorld task in which the agent settles a shared dinner bill among roommates. Information from the contact result produced at Turn 8 is used by later operations at Turns 9, 15, and 16. At Turn 18, the agent needs Nancy’s email address, which has not been copied forward and remains available only from the Turn 8 result.
Figure 3: ContextRender architecture. The runtime adapter records messages, tool calls, and complete tool results in a persistent context graph. Tool-Flow Analysis records observed reuse as execution dependencies. Before each model invocation, the renderer includes the accumulated execution history if it fits within the history budget B ; otherwise, it selects tool results using observed reuse, recency, and semantic relevance. Omitted results remain stored for later invocations.
Figure 4: Overall performance (AppWorld TGC; QA F1) across all tasks versus mean standardized inference cost on the same tasks that trigger context management at B=6 K. Costs are normalized to the mean cost of Full history ( =1 ). “Recency” denotes Recency + Summary.
Model
Benchmark
Full history
Prune
Compact
Recency + Summary
Semantic Retrieval
ACON
Perf.
Cost
Perf.
Cost
Perf.
Cost
Perf.
Cost
Perf.
Cost
Perf.
Cost
GPT-4.1
AppWorld Normal
↓ 2.3%
↓ 32.2%
↑ 14.4%
↓ 3.2%
↑ 12.4%
↓ 0.8%
↑ 7.6%
↓ 1.6%
↑ 7.6%
↓ 18.4%
↑ 3.3%
↓ 14.9%
AppWorld Challenge
↑ 8.5%
↓ 28.8%
↑ 19.7%
↓ 2.7%
↑ 12.1%
↓ 1.8%
↑ 11.6%
↓ 16.7%
↑ 14.4%
↓ 9.7%
↑ 11.1%
↓ 8.1%
8-objective QA
↑ 4.5%
↓ 10.2%
↑ 3.7%
↓ 2.0%
↑ 2.1%
↓ 4.5%
↑ 9.5%
→ 0.0%
↑ 14.4%
→ 0.0%
↑ 10.4%
↓ 5.7%
Muse Glimmer 30B
AppWorld Normal
↑ 5.8%
↓ 12.4%
↑ 14.2%
↑ 1.7%
↑ 7.4%
↑ 0.8%
↑ 8.2%
↓ 3.2%
↑ 10.7%
↓ 11.1%
↑ 9.8%
↓ 2.4%
AppWorld Challenge
↑ 2.9%
↓ 19.9%
↑ 16.8%
↓ 7.2%
↑ 9.2%
↓ 16.8%
↑ 10.1%
↓ 2.3%
↑ 13.6%
↓ 8.5%
↑ 8.2%
↓ 12.2%
Table 1: Relative changes in task performance and mean inference cost of ContextRender versus each baseline and Full history . Green, red, and gray arrows indicate improvements, regressions, and no change, respectively.
Variant
Performance
Later-Used Result Recall
AppWorld Later-Used Result Recall by Result Age
AppWorld (TGC)
QA (F1)
AppWorld
QA
≤3
4–10
11–25
>25
CR–Recency
71.4
48.5
63.8
86.0
95.7
78.7
37.7
18.5
CR–Relevance
73.2
47.5
71.4
64.2
89.5
81.9
54.1
47.1
CR–Recency + Relevance
72.6
50.8
61.6
76.1
84.4
72.5
43.9
22.7
CR–Full (+Reuse)
75.6
53.1
87.2
91.0
99.3
91.5
79.0
68.5
Table 2: Ablation of ContextRender with GPT-4.1 on AppWorld Test-Normal and 8-objective QA at B=6 K. Recall denotes Later-Used Result Recall. Result age is the number of turns between a tool result’s creation and its subsequent reuse. All values are percentages.
Figure 5: Budget sensitivity with GPT-4.1. B denotes the history budget. “Recency” denotes Recency + Summary.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Selector
Wrec
Wrel
Wreuse
CR–Recency
1.0
0.0
0.0
CR–Relevance
0.0
1.0
0.0
CR–Recency + Relevance
1.0
1.0
0.0
CR–Full (+Reuse)
1.0
1.0
1.0
Appendix
Table 3: Scoring weights used in the selector ablations. Zero disables the corresponding term; retained weights remain unchanged.
Result age at reuse
Events
≤3 turns
713
4–10 turns
1,124
11–25 turns
1,099
>25 turns
238
Total
3,174
Appendix
Table 4: AppWorld reuse events by result age.
Result age at reuse
Events
≤3 turns
224
4–10 turns
70
11–25 turns
41
>25 turns
0
Total
335
Appendix
Table 5: 8-objective QA reuse events by result age.
Method
AppWorld Test-Normal
AppWorld Test-Challenge
8-objective QA
TGC (%)
Cost
Norm.
TGC (%)
Cost
Norm.
F1 (%)
Cost
Norm.
Full history
77.4
0.177
1.000
51.1
0.302
1.000
50.8
0.167
1.000
Prune
66.1
0.124
0.701
46.3
0.221
0.732
51.2
0.153
0.916
Compact (LLM)
67.3
0.121
0.684
49.4
0.219
0.725
52.0
0.157
0.940
Recency + Summary
70.2
0.122
0.689
49.6
0.258
0.854
48.5
0.150
0.898
Semantic Retrieval
70.2
0.147
0.831
48.4
0.238
0.788
46.4
0.150
0.898
Appendix
Table 6: Overall task performance and mean inference cost for GPT-4.1. Performance is measured across all tasks; cost is measured on the same tasks that trigger context management when accumulated execution history exceeds B=6 K. Cost is in reference USD; normalized cost uses uncapped Full history ( =1 ) on the same tasks in each benchmark.
Method
AppWorld Test-Normal
AppWorld Test-Challenge
8-objective QA
TGC (%)
Cost
Norm.
TGC (%)
Cost
Norm.
F1 (%)
Cost
Norm.
Full history
81.5
0.137
1.000
58.3
0.161
1.000
50.4
0.269
1.000
Prune
75.6
0.118
0.861
51.3
0.139
0.863
49.3
0.243
0.903
Compact (LLM)
80.4
0.119
0.869
54.9
0.155
0.963
50.4
0.259
0.963
Recency + Summary
79.8
0.124
0.905
54.4
0.132
0.820
48.2
0.245
0.911
Semantic Retrieval
78.0
0.135
0.985
52.8
0.141
0.876
51.1
0.235
0.874
Appendix
Table 7: Overall task performance and mean inference cost for Muse Glimmer 30B. Performance is measured across all tasks; cost is measured on the same tasks that trigger context management when accumulated execution history exceeds B=6 K. Cost is in reference USD; normalized cost uses uncapped Full history ( =1 ) on the same tasks in each benchmark.
Method
AppWorld Test-Normal
AppWorld Test-Challenge
8-objective QA
TGC (%)
Cost
Norm.
TGC (%)
Cost
Norm.
F1 (%)
Cost
Norm.
Full history
82.7
0.1425
1.000
77.0
0.236
1.000
56.5
0.476
1.000
Prune
81.0
0.1185
0.832
57.6
0.180
0.763
55.2
0.424
0.891
Compact (LLM)
82.1
0.126
0.884
70.7
0.222
0.941
58.2
0.444
0.933
Recency + Summary
82.1
0.117
0.821
71.2
0.179
0.758
58.5
0.435
0.914
Semantic Retrieval
78.0
0.120
0.842
67.1
0.162
0.686
55.3
0.419
0.880
Appendix
Table 8: Overall task performance and mean inference cost for DeepSeek V4 Pro. Performance is measured across all tasks; cost is measured on the same tasks that trigger context management when accumulated execution history exceeds B=6 K. Cost is in reference USD; normalized cost uses uncapped Full history ( =1 ) on the same tasks in each benchmark.
Agent runtime
Insertion point
Integration changes
Used for
opencode
experimental.chat.messages.transform plugin hook; the plugin is specified in opencode.json
The runtime’s own compaction and pruning are disabled so that only ContextRender manages historical context.
Integration only
AppWorld ReAct agent
trimmed_messages property replaced by a mixin on the live agent object
One property. Prompting, parsing, retry logic, tool use, runner, configuration, and CLI are unchanged.
AppWorld
ACON UnifiedAgent (smolagents)
MemoryManager subclass overriding get_conversation_history ; the class reference is rebound before agent construction
Memory class only. Runner, agent, prompts, retriever, and grading are unchanged.
8-objective QA
OpenClaw (CLI in Docker)
OpenAI-compatible proxy registered as a custom provider in openclaw.json
No runtime code is modified; context rendering is applied by the proxy before the request is forwarded.
Sanity check
Meta ARE default agent (GAIA2)
The same proxy, selected with the --provider local --endpoint flags
No runtime code is modified.
Integration check
Appendix
Table 9: Agent runtimes connected to ContextRender . For each runtime, we list the adapter insertion point, the runtime component modified by the integration, and its use in this work. Prompting, tool execution, parsing, and grading are not rewritten.