At the start of every session, LLM agents load a fixed context file, such as AGENTS.md. Each loaded token in the file is charged again in every later round of the session, and these files can degrade performance as they grow in size. However, in practice, human or automated curators usually grow these files by appending. We formulate context curation as a capacitated assortment problem. Instructions consume tokens under a finite attention capacity; adding an instruction never raises the compliance of the others, while retained instructions incur a per-session setup cost. We prove an upper bound on the optimal file size, regardless of the number of available candidate instructions, and that appending every instruction with positive standalone value can be arbitrarily worse in net value than selecting an optimal subset. A token budget also limits the loss when the token price is underestimated. We then examine what can be learned from past sessions and how this information can guide decisions to add or remove instructions. Feedback is inherently censored: the benefits and harms of loaded instructions are observable, whereas missing instructions generate feedback only when their absence causes harm. In this setting, we show that deleting instructions ignored by agents can inevitably remove helpful ones. We characterize how much evidence should be collected before adding an instruction. Besides, we bound regret when human reviewers can inspect only a limited number of edits per period. Empirical experiments further show that irrelevant rules drawn from real context files reduce language-model compliance.
Figures & tables
Figure 1: Gross value against file size for Claude Haiku 4.5 and Claude Sonnet 5.5. Gross value counts relevant instructions that the model follows, after adjustment for chance keyword matches. Each panel uses a different candidate pool. Points show means across 60 tasks. Bars show 95% confidence intervals from a bootstrap that resamples tasks and keeps each task’s runs together. The last point in each panel loads the entire pool.
Figure 2: Regret from the edit limit against N/m . Both axes use logarithmic scales. Each panel uses a different epoch length L . Each point combines simulation settings with the same N/m . The lines show the lower bound from Theorem 3 and the cap term of the upper bound from Theorem 2 . The variation across settings with the same N/m is smaller than the plotted marker.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Dilution
q(n)
M(n)
nˉ
Single choice
1+nuu
<1
⌊1/ρ−1/u⌋+
Exponential
αe−bn
≤beα
⌊ln(α/ρ)/b⌋+
Power ( β<1 )
αn−β
αn1−β
⌊(α/ρ)1/β⌋
None
α
αn
∞ if α≥ρ , else 0
Appendix
Table 1: Identical instructions with compliance q(n) in a file of n , where 0<α≤1 , b,u>0 , 0<β<1 , u=v/v0 and ⌊y⌋+=max(0,⌊y⌋) . Assumption 2 holds in the first three rows for every ρ , and in the last iff α<ρ . The nˉ column assumes wˉ>0 ; if wˉ=0 , then N={0} and nˉ=0 .
Figure 3: Size sweep: value Rρexp(Sb) at price ρexp=0.1 against file size b for both models, one panel per pool size N . Points are means over the 60 tasks; bars are ±1.96 Newey–West standard errors over the task sequence. The stars mark the measured optima n⋆(N) .
Figure 4: Size sweep: marginal relevant compliance MC(b) against relevant load at N=3200 for both models, with 95% task-cluster bootstrap bars.
N
b
R0(Sb)
raw
ℓ(Sb)
MC (step to b )
200
5
0.53
0.60
1.33
10
1.19
1.34
1.58
2.67 [1.91, 3.88]
20
2.05
2.35
3.39
0.48 [0.37, 0.57]
40
3.18
3.89
8.69
0.21 [0.13, 0.28]
80
5.11
6.49
14.00
0.36 [0.23, 0.50]
200
3.84
6.25
25.06
−0.11 [ −0.22 , −0.01 ]
Appendix
Table 2: Size-sweep cells for Claude Haiku 4.5: gross value R0(Sb) (followed relevant instructions per task, net of chance), the raw count, the singleton-weighted relevant load ℓ(Sb) , and the marginal compliance MC of the step ending at b with its 95% task-cluster bootstrap interval.
N
b
R0(Sb)
raw
ℓ(Sb)
MC (step to b )
200
5
0.97
1.04
1.56
10
2.25
2.39
3.19
0.78 [0.67, 0.88]
20
2.27
2.56
5.28
0.01 [ −0.09 , 0.12]
40
8.07
8.74
11.28
0.97 [0.83, 1.11]
80
10.24
11.55
19.00
0.28 [0.05, 0.51]
200
17.21
19.45
35.22
0.43 [0.25, 0.60]
Appendix
Table 3: Size-sweep cells for Claude Sonnet 5.5, as in Table 2 .
N
ρexp=0.05
ρexp=0.1
ρexp=0.2
ρexp=0.4
Haiku
200
40; 6.9 [5.9, 8.0]
10; 15.8 [14.6, 16.8]
5; 34.5 [33.3, 35.5]
5; 72.3 [71.1, 73.3]
800
40; 37.5 [36.4, 38.7]
5; 76.2 [74.9, 77.3]
5; 153.7 [152.4, 154.8]
5; 308.7 [307.4, 309.8]
3200
20; 159.1 [158.4, 159.9]
10; 316.7 [316.0, 317.4]
5; 632.2 [631.5, 632.9]
5; 1263.5 [1262.8, 1264.1]
Sonnet
200
200; 0.0 [0.0, 1.3]
40; 6.0 [3.7, 8.3]
10; 21.9 [19.3, 24.5]
5; 59.3 [56.7, 62.0]
800
400; 41.9 [37.9, 47.6]
20; 75.9 [73.1, 78.4]
5; 152.3 [149.5, 154.6]
5; 307.3 [304.4, 309.6]
3200
200; 166.9 [161.3, 173.0]
40; 320.4 [315.1, 324.2]
40; 632.1 [626.7, 635.9]
5; 1260.2 [1254.7, 1263.7]
Appendix
Table 4: Size sweep: measured optimum n⋆(N) and accretion loss Rρexp(Sn⋆)−Rρexp(SN) with 95% task-cluster bootstrap interval, per model, pool size and price ρexp .
Figure 5: Size sweep, Claude Haiku 4.5: value Rρexp(Sb) against file size at each price ρexp , one series per pool size N , with Newey–West error bars and stars at n⋆(N) as in Figure 3 .
Figure 6: Size sweep, Claude Sonnet 5.5: value Rρexp(Sb) against file size at each price ρexp , as in Figure 5 .
Figure 7: Monotone dilution, item by item, Claude Haiku 4.5: compliance of the relevant items of the b=40 file in that file and in every larger file of the same pool, with 95% task-cluster bootstrap bars.
Figure 8: Monotone dilution, item by item, Claude Sonnet 5.5, as in Figure 7 .
configuration
K
class
best form
M(500)/500
llama-4-scout
34
weak
power
0.067
claude-3.5-haiku
52
weak
hyperbola
0.085
gpt-4o-mini
61
weak
hyperbola
0.104
gpt-4o
81
weak
hyperbola
0.154
llama-4-maverick
87
weak
power
0.096
gpt-4.1-nano
99
weak
hyperbola
0.162
Appendix
Table 5: Dilution: fitted capacity K , class, AIC-best curve form and compliance M(500)/500 at n=500 , per configuration.
Figure 9: Dilution: q(n)=M(n)/n against n for the lowest-, median- and highest- K configuration of each class (three left panels), and the 24 late/early slope ratios sorted, with 95% seed-bootstrap bars and the 0.1 reference line (right).
N
m=2
m=4
m=8
m=16
16
4.35 (7.77)
4.83 (8.05)
N<4m
N<4m
32
4.20 (6.18)
4.42 (6.33)
5.02 (6.61)
N<4m
64
4.13 (5.39)
4.26 (5.46)
4.57 (5.61)
5.38 (5.90)
128
4.09 (5.00)
4.19 (5.03)
4.40 (5.10)
4.87 (5.25)
256
4.08 (4.80)
4.15 (4.82)
4.32 (4.85)
4.66 (4.93)
Appendix
Table 6: Regret experiment: cap-term regret of ETC- m over the lower bound (ΔL/2)(nm−1) , with the plan’s upper limit 4(Δmax/Δ)(1+m(1+nB/s)/(2N)) in parentheses. The ratio does not depend on L . Cells with N<4m are outside the call.