For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the exact originals retrievable, but editing the history can break prefix-cache reuse, and prior recoverable methods time their edits by forecasts of future reuse or by preset intervals. We propose CADOC (Cache-Aware Dynamic Object Context), an online algorithm that replaces structured objects with compact Cards while preserving exact, on-demand retrieval of their original contents. CADOC schedules replacements in batches by balancing accumulated waiting cost against shared cache-reconstruction cost. Its scheduling rule follows from an economic order quantity trade-off, recovers the optimal integer batch under stationary assumptions. Across evaluation, CADOC consistently achieves the lowest aggregate input cost among the compared configurations, which reduces input cost by approximately 40% on average while maintaining task performance close to full context. CADOC thus provides a cost-derived approach to compressible context management, demonstrating that efficient compression depends not only on shortening prompts but also on scheduling edits to preserve cache reuse.
Figures & tables
Figure 1: (a) Immediate cuts cumulative input tokens by 55.93% but raises input cost by 3.67% versus Full Context. (b) Pending objects become Cards in place; the hot tail stays visible, and Working Memory returns exact contents.
Figure 2: The cached interval and the batch-commit rule. Left: the b pending blocks shorten to b(x−g) tokens, while the h−1 cached hot blocks stay raw. Right: waiting-loss updates and the commit decision.
Figure 3: Aggregate input-cost savings relative to Full Context over HCE. Negative values indicate costs above Full Context.
Figure 4: Commit costs and batch sizes across hot-tail settings. (a) Aggregate input costs of fixed-batch configurations and CADOC, normalized by Full Context. (b) CADOC’s mean number of blocks per committed batch at h=2,4,8,16 .
Benchmark
FULL solved
CADOC solved
Input ↓
Cost ↓
Terminal-Bench 2.1 ( Merrill et al., 2026 )
38/89
36/89
49.40%
47.13%
LongMemEval-V2 ( Wu et al., 2026a )
10/18
10/18
80.60%
42.40%
SWE-bench ( Jimenez et al., 2024 )
31/50
29/50
44.72%
34.12%
Table 1: Complete benchmark-chain results. Input denotes cumulative input tokens, Cost uses N+0.1R , and both reductions are relative to FULL.
Figure 5: Cumulative input cost during the three benchmark chains. The horizontal axis counts completed tasks. Gray dashed curves denote FULL and blue curves denote CADOC.
Figure 6: Input-cost savings from offline replay of the conversation histories retained by FULL on three benchmarks.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
h=2
h=4
h=8
h=16
Policy
Passes
Macro
Passes
Macro
Passes
Macro
Passes
Macro
Full
57/93
4.318
57/93
4.318
57/93
4.318
57/93
4.318
Immediate
64/93
4.326
59/93
4.274
64/93
4.285
63/93
4.404
Fixed b=4
65/93
4.318
64/93
4.346
62/93
4.305
61/93
4.354
Fixed b=8
63/93
4.336
59/93
4.263
64/93
4.353
69/93
4.538
Token 16,384
65/93
4.421
60/93
4.247
67/93
4.492
67/93
4.447
Appendix
Table 2: Continuation passes and session-macro scores for six policies across four hot-tail settings.
Boundary
W
G
Q
W+0.1G
Decision
14
3,424.9
13,622
5,688.9
4,787.1
Wait
15
4,787.1
14,878
6,124.5
6,274.9
Commit
Appendix
Table 3: Two decisions to show the process; boundaries are zero-based request indices.
Figure 7: Cache-read weight sensitivity. Top: savings relative to Full for CADOC and the best of nine baselines, including Full, at each weight. Bottom: savings relative to the best of eight other compression policies, with separate vertical scales starting at zero. Markers are measured points; lines guide the eye.
wr
Terminal-Bench
SWE-bench
LongMemEval
Pooled
0.05
45.73%
40.48%
62.52%
45.06%
0.1
49.83%
46.75%
75.57%
50.39%
0.25
52.66%
51.61%
86.28%
54.26%
0.5
53.71%
53.57%
90.55%
55.75%
1
54.12%
54.72%
92.85%
56.48%
Appendix
Table 4: CADOC input-cost savings relative to Full at five cache-read weights. Pooled savings use the ratio of summed costs across the three benchmark histories.