At the start of every session, LLM agents load a fixed context file, such as AGENTS.md. Each loaded token in the file is charged again in every later round of the session, and these files can degrade performance as they grow in size. However, in practice, human or automated curators usually grow these files by appending. We formulate context curation as a capacitated assortment problem. Instructions consume tokens under a finite attention capacity; adding an instruction never raises the compliance of the others, while retained instructions incur a per-session setup cost. We prove an upper bound on the optimal file size, regardless of the number of available candidate instructions, and that appending every instruction with positive standalone value can be arbitrarily worse in net value than selecting an optimal subset. A token budget also limits the loss when the token price is underestimated. We then examine what can be learned from past sessions and how this information can guide decisions to add or remove instructions. Feedback is inherently censored: the benefits and harms of loaded instructions are observable, whereas missing instructions generate feedback only when their absence causes harm. In this setting, we show that deleting instructions ignored by agents can inevitably remove helpful ones. We characterize how much evidence should be collected before adding an instruction. Besides, we bound regret when human reviewers can inspect only a limited number of edits per period. Empirical experiments further show that irrelevant rules drawn from real context files reduce language-model compliance.
Figures & tables
Figure 1: Gross value against file size for Claude Haiku 4.5 and Claude Sonnet 5.5. Gross value counts relevant instructions that the model follows, after adjustment for chance keyword matches. Each panel uses a different candidate pool. Points show means across 60 tasks. Bars show 95% confidence intervals from a bootstrap that resamples tasks and keeps each task’s runs together. The last point in each panel loads the entire pool.
Figure 2: Regret from the edit limit against N/m . Both axes use logarithmic scales. Each panel uses a different epoch length L . Each point combines simulation settings with the same N/m . The lines show the lower bound from Theorem 3 and the cap term of the upper bound from Theorem 2 . The variation across settings with the same N/m is smaller than the plotted marker.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Dilution
q(n)
M(n)
nˉ
Single choice
1+nuu
<1
⌊1/ρ−1/u⌋+
Exponential
αe−bn
≤beα
⌊ln(α/ρ)/b⌋+
Power ( β<1 )
αn−β
αn1−β
⌊(α/ρ)1/β⌋
None
α
αn
∞ if α≥ρ , else 0
Appendix
Table 1: Identical instructions with compliance q(n) in a file of n , where 0<α≤1 , b,u>0 , 0<β<1 , u=v/v0 and ⌊y⌋+=max(0,⌊y⌋) . Assumption 2 holds in the first three rows for every ρ , and in the last iff α<ρ . The nˉ column assumes wˉ>0 ; if wˉ=0 , then N={0} and nˉ=0 .
Figure 3: Size sweep: value Rρexp(Sb) at price ρexp=0.1 against file size b for both models, one panel per pool size N . Points are means over the 60 tasks; bars are ±1.96 Newey–West standard errors over the task sequence. The stars mark the measured optima n⋆(N) .
Figure 4: Size sweep: marginal relevant compliance MC(b) against relevant load at N=3200 for both models, with 95% task-cluster bootstrap bars.
N
b
R0(Sb)
raw
ℓ(Sb)
MC (step to b )
200
5
0.53
0.60
1.33
10
1.19
1.34
1.58
2.67 [1.91, 3.88]
20
2.05
2.35
3.39
0.48 [0.37, 0.57]
40
3.18
3.89
8.69
0.21 [0.13, 0.28]
80
5.11
6.49
14.00
0.36 [0.23, 0.50]
200
3.84
6.25
25.06
−0.11 [ −0.22 , −0.01 ]
Appendix
Table 2: Size-sweep cells for Claude Haiku 4.5: gross value R0(Sb) (followed relevant instructions per task, net of chance), the raw count, the singleton-weighted relevant load ℓ(Sb) , and the marginal compliance MC of the step ending at b with its 95% task-cluster bootstrap interval.
N
b
R0(Sb)
raw
ℓ(Sb)
MC (step to b )
200
5
0.97
1.04
1.56
10
2.25
2.39
3.19
0.78 [0.67, 0.88]
20
2.27
2.56
5.28
0.01 [ −0.09 , 0.12]
40
8.07
8.74
11.28
0.97 [0.83, 1.11]
80
10.24
11.55
19.00
0.28 [0.05, 0.51]
200
17.21
19.45
35.22
0.43 [0.25, 0.60]
Appendix
Table 3: Size-sweep cells for Claude Sonnet 5.5, as in Table 2 .
N
ρexp=0.05
ρexp=0.1
ρexp=0.2
ρexp=0.4
Haiku
200
40; 6.9 [5.9, 8.0]
10; 15.8 [14.6, 16.8]
5; 34.5 [33.3, 35.5]
5; 72.3 [71.1, 73.3]
800
40; 37.5 [36.4, 38.7]
5; 76.2 [74.9, 77.3]
5; 153.7 [152.4, 154.8]
5; 308.7 [307.4, 309.8]
3200
20; 159.1 [158.4, 159.9]
10; 316.7 [316.0, 317.4]
5; 632.2 [631.5, 632.9]
5; 1263.5 [1262.8, 1264.1]
Sonnet
200
200; 0.0 [0.0, 1.3]
40; 6.0 [3.7, 8.3]
10; 21.9 [19.3, 24.5]
5; 59.3 [56.7, 62.0]
800
400; 41.9 [37.9, 47.6]
20; 75.9 [73.1, 78.4]
5; 152.3 [149.5, 154.6]
5; 307.3 [304.4, 309.6]
3200
200; 166.9 [161.3, 173.0]
40; 320.4 [315.1, 324.2]
40; 632.1 [626.7, 635.9]
5; 1260.2 [1254.7, 1263.7]
Appendix
Table 4: Size sweep: measured optimum n⋆(N) and accretion loss Rρexp(Sn⋆)−Rρexp(SN) with 95% task-cluster bootstrap interval, per model, pool size and price ρexp .
Figure 5: Size sweep, Claude Haiku 4.5: value Rρexp(Sb) against file size at each price ρexp , one series per pool size N , with Newey–West error bars and stars at n⋆(N) as in Figure 3 .
Figure 6: Size sweep, Claude Sonnet 5.5: value Rρexp(Sb) against file size at each price ρexp , as in Figure 5 .
Figure 7: Monotone dilution, item by item, Claude Haiku 4.5: compliance of the relevant items of the b=40 file in that file and in every larger file of the same pool, with 95% task-cluster bootstrap bars.
Figure 8: Monotone dilution, item by item, Claude Sonnet 5.5, as in Figure 7 .
configuration
K
class
best form
M(500)/500
llama-4-scout
34
weak
power
0.067
claude-3.5-haiku
52
weak
hyperbola
0.085
gpt-4o-mini
61
weak
hyperbola
0.104
gpt-4o
81
weak
hyperbola
0.154
llama-4-maverick
87
weak
power
0.096
gpt-4.1-nano
99
weak
hyperbola
0.162
Appendix
Table 5: Dilution: fitted capacity K , class, AIC-best curve form and compliance M(500)/500 at n=500 , per configuration.
Figure 9: Dilution: q(n)=M(n)/n against n for the lowest-, median- and highest- K configuration of each class (three left panels), and the 24 late/early slope ratios sorted, with 95% seed-bootstrap bars and the 0.1 reference line (right).
N
m=2
m=4
m=8
m=16
16
4.35 (7.77)
4.83 (8.05)
N<4m
N<4m
32
4.20 (6.18)
4.42 (6.33)
5.02 (6.61)
N<4m
64
4.13 (5.39)
4.26 (5.46)
4.57 (5.61)
5.38 (5.90)
128
4.09 (5.00)
4.19 (5.03)
4.40 (5.10)
4.87 (5.25)
256
4.08 (4.80)
4.15 (4.82)
4.32 (4.85)
4.66 (4.93)
Appendix
Table 6: Regret experiment: cap-term regret of ETC- m over the lower bound (ΔL/2)(nm−1) , with the plan’s upper limit 4(Δmax/Δ)(1+m(1+nB/s)/(2N)) in parentheses. The ratio does not depend on L . Cells with N<4m are outside the call.
LLM agents increasingly face long-horizon tasks such as web search and deep research in real-world applications, where accumulated context can cause long-context degradation and reasoning failures. Prior work mitigates this through context management with agent-side context control or fixed strategies such as summarization, which require training the agent itself for adaptation - making it impractical for closed-source agents and ignoring that different agents may require different strategies. We introduce Adaptive Context Management (AdaCoM), which trains an external LLM to manage the context of a frozen agent through flexible modification actions and end-to-end reinforcement learning. Across diverse agents on web search and deep research benchmarks, AdaCoM substantially improves performance by preserving task constraints and progress while pruning stale content. The learned strategies reveal a Fidelity-Reliability Trade-off: agents with higher vanilla ReAct performance benefit from higher-fidelity context preservation, whereas lower-performing agents require more aggressive compression to stay within a reliable reasoning regime. Transfer experiments show that AdaCoM generalizes most effectively across agents with similar capability (measured by vanilla ReAct performance), suggesting a practical path toward reusable context managers for agent systems.
Lu Yi, Runlin Lei, Liuyi Yao +6
1Renmin University of China · Work done during internship at Tongyi Lab, Alibaba Group · 2Tongyi Lab, Alibaba Group +2
Large Language Models (LLMs) have a bounded context window. The context window is the maximum input size an LLM can consume for a single inference. AI agents rely on a process called context compaction to fit their state within the context window when calling an LLM. Despite its ubiquity, context compaction has received essentially no formal analysis. In this paper, we initiate a formal study of context compaction. We first introduce a framework consisting of two games that capture the two algorithmic strategies for context compaction used by contemporary AI agents in practice. The Context Selection Game models context compaction algorithms that select a subset of an agent's accumulated state to retain. The Context Generation Game models context compaction algorithms that summarize an agent's state by an arbitrary message of bounded length. We then prove an equivalence between the Context Generation Game and one-way communication complexity. The minimum context compaction budget for answering a set of queries within a target error is equal to the one-way communication complexity of the induced communication problem at the same error. Known bounds from communication complexity therefore transfer directly to context compaction. We also show that the Context Selection Game corresponds to a restricted class of one-way communication protocols. Any gap between selection and generation is therefore a gap between two classes of communication protocols. We prove that there exists a set of queries for which generation needs strictly less budget than selection. The equivalence between the Context Generation Game and one-way communication also lets us measure how well a deployed context compaction algorithm performs on a query relative to the optimal strategy. As an example, we present a case study that evaluates Anthropic's context compaction endpoint on set membership queries.
Hayder Tirmazi, Sam Markelon, Allison Bishop +1
AllSpice Inc. · Boston University · Proof Trading +2
Modern large language model (LLM) agents do not simply need longer contexts; they need decision-relevant evidence at the moment of action. We study decision-aware context selection: ranking retrieved files, tests, traces, rules, and memories by their expected effect on an agent's next action rather than by semantic similarity alone. We present the Counterfactual-Inspired Context Layer (CICL), which builds an instance context graph, estimates decision-oriented utility for candidate units, and compresses selected evidence into typed memory cards. The same schema can be instantiated with hosted LLM judges, local surrogates, or lightweight rankers, making the selection protocol auditable across model choices. On 50 SWE-bench Verified file-retrieval instances, Qwen3.6-Plus reranking of BM25 top-50 candidates improves hit@1 from 0.58 to 0.78 and MRR@10 from 0.634 to 0.790, with all 2,500 judgments parseable. Controlled diagnostics show that CICL identifies action-critical evidence: removing the top-utility semantic unit reduces F1 from 0.245 to 0.000. In selected-then-compressed mode, memory cards save 44.93 tokens per query while preserving selected evidence. CICL provides a practical layer for measuring, ranking, and compressing decision-critical context for tool-using agents. Code is available at https://github.com/stephen-guan-researcher/CICL.