For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the exact originals retrievable, but editing the history can break prefix-cache reuse, and prior recoverable methods time their edits by forecasts of future reuse or by preset intervals. We propose CADOC (Cache-Aware Dynamic Object Context), an online algorithm that replaces structured objects with compact Cards while preserving exact, on-demand retrieval of their original contents. CADOC schedules replacements in batches by balancing accumulated waiting cost against shared cache-reconstruction cost. Its scheduling rule follows from an economic order quantity trade-off, recovers the optimal integer batch under stationary assumptions. Across evaluation, CADOC consistently achieves the lowest aggregate input cost among the compared configurations, which reduces input cost by approximately 40% on average while maintaining task performance close to full context. CADOC thus provides a cost-derived approach to compressible context management, demonstrating that efficient compression depends not only on shortening prompts but also on scheduling edits to preserve cache reuse.
Figures & tables
Figure 1: (a) Immediate cuts cumulative input tokens by 55.93% but raises input cost by 3.67% versus Full Context. (b) Pending objects become Cards in place; the hot tail stays visible, and Working Memory returns exact contents.
Figure 2: The cached interval and the batch-commit rule. Left: the b pending blocks shorten to b(x−g) tokens, while the h−1 cached hot blocks stay raw. Right: waiting-loss updates and the commit decision.
Figure 3: Aggregate input-cost savings relative to Full Context over HCE. Negative values indicate costs above Full Context.
Figure 4: Commit costs and batch sizes across hot-tail settings. (a) Aggregate input costs of fixed-batch configurations and CADOC, normalized by Full Context. (b) CADOC’s mean number of blocks per committed batch at h=2,4,8,16 .
Benchmark
FULL solved
CADOC solved
Input ↓
Cost ↓
Terminal-Bench 2.1 ( Merrill et al., 2026 )
38/89
36/89
49.40%
47.13%
LongMemEval-V2 ( Wu et al., 2026a )
10/18
10/18
80.60%
42.40%
SWE-bench ( Jimenez et al., 2024 )
31/50
29/50
44.72%
34.12%
Table 1: Complete benchmark-chain results. Input denotes cumulative input tokens, Cost uses N+0.1R , and both reductions are relative to FULL.
Figure 5: Cumulative input cost during the three benchmark chains. The horizontal axis counts completed tasks. Gray dashed curves denote FULL and blue curves denote CADOC.
Figure 6: Input-cost savings from offline replay of the conversation histories retained by FULL on three benchmarks.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
h=2
h=4
h=8
h=16
Policy
Passes
Macro
Passes
Macro
Passes
Macro
Passes
Macro
Full
57/93
4.318
57/93
4.318
57/93
4.318
57/93
4.318
Immediate
64/93
4.326
59/93
4.274
64/93
4.285
63/93
4.404
Fixed b=4
65/93
4.318
64/93
4.346
62/93
4.305
61/93
4.354
Fixed b=8
63/93
4.336
59/93
4.263
64/93
4.353
69/93
4.538
Token 16,384
65/93
4.421
60/93
4.247
67/93
4.492
67/93
4.447
Appendix
Table 2: Continuation passes and session-macro scores for six policies across four hot-tail settings.
Boundary
W
G
Q
W+0.1G
Decision
14
3,424.9
13,622
5,688.9
4,787.1
Wait
15
4,787.1
14,878
6,124.5
6,274.9
Commit
Appendix
Table 3: Two decisions to show the process; boundaries are zero-based request indices.
Figure 7: Cache-read weight sensitivity. Top: savings relative to Full for CADOC and the best of nine baselines, including Full, at each weight. Bottom: savings relative to the best of eight other compression policies, with separate vertical scales starting at zero. Markers are measured points; lines guide the eye.
wr
Terminal-Bench
SWE-bench
LongMemEval
Pooled
0.05
45.73%
40.48%
62.52%
45.06%
0.1
49.83%
46.75%
75.57%
50.39%
0.25
52.66%
51.61%
86.28%
54.26%
0.5
53.71%
53.57%
90.55%
55.75%
1
54.12%
54.72%
92.85%
56.48%
Appendix
Table 4: CADOC input-cost savings relative to Full at five cache-read weights. Pooled savings use the ratio of summed costs across the three benchmark histories.
Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existing context compression methods inevitably incur information loss and are triggered by rigid heuristic rules, leaving them misaligned with the agent's evolving reasoning focus. We propose Agentic Context Management (ACM), a framework that equips agents with purpose-built context editing tools for lossless context management. Inspired by the interaction between short-term and long-term human memory, the agent autonomously decides when to compress its context, offloads discarded content to an external memory system, and queries it on demand for later retrieval. Building on this framework, we further develop a post-training pipeline that constructs high-quality demonstrations of context management and improves model performance on both agentic search and coding tasks. Further analysis reveals that effective context management reduces peak token pressure, enables extended explorations, and yields more consistent solutions across independent trials. Code, data, and model checkpoints are available at https://github.com/lixiaochuan2020/agentic-context-management.
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing is too narrow. Actively managing what an agent holds in mind is a lifecycle, not merely a store: it spans deciding what to remember, extracting and structuring it, choosing the right store per data type, consolidating and forgetting while preserving provenance, deciding what is relevant now, anticipating what is needed next, and compacting context to a budget without losing what matters. In serious production this operates not over a single user but across an organizational scope hierarchy. We name this discipline Agentic Context Management (ACM) and decompose it into five primitives: architecting, ingesting, scoping, anticipating, and compacting & consolidation. We then make the economic case: naive context accumulation grows token cost quadratically in conversation length, crude summarization buys linear cost at the price of an accuracy cliff, and only validated compaction achieves linear cost with preserved fidelity. We describe a reference implementation, Maximem Synap, that realizes the five primitives as a multi-tenant service and reports 92% on LongMemEval and 93.2% on LoCoMo under the configuration detailed in Section 6. We close with dimensions existing benchmarks do not yet capture, latency, token efficiency, and context-rot resistance, and the frontier of decision-level and organization-level context the category points toward.
Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench. The per-rollout savings of CliffCompaction make the performance--cost trade-off of test-time scaling more efficient, adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it. We never compact a compaction---each pass operates only on original content, and prior compacted output is discarded, preventing context drift from accumulating. These properties sustain continual learning over sessions exceeding a million tokens: on KernelBench, CliffCompaction reaches CUDA kernel speedups of 2.23× after 200 steps and 3.58× after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique. We open-source a scaffold-agnostic API-proxy implementation of CliffCompaction usable with Claude Code, Codex and other harnesses.