Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.
Figures & tables
Figure 1: ReFold’s savings carry from one session to the serving system. With Qwen3.6-27B on SWE-bench Verified, it cuts (a) tokens and (b) peak KV cache per session, so a vLLM server (c) fits 2× the sessions in the same memory and (d) runs tasks up to 3.4× cheaper and 1.7× faster.
Figure 2: The six strategies on the same 50 SWE-bench Verified tasks at a 16 k budget: per-session KV-cache footprint and peak composition (a, b), per-task cost and duration with resolve rate (c, d), prefix-cache reuse along one trajectory (e) and the fate of evicted observation tokens (f).
Figure 3: ReFold compresses the rendered context while preserving the append-only history. Turn 3 repeats turn 1 and becomes a stub, and turns 4–5, reported finished, collapse into a note. Turns 7–9 , past the k=3 boundary, render in full. Stubs and notes can be restored from the history.
Resolve
Steps
Tokens (k)
KV cache (GiB)
Cost (¢)
Benchmark
base
ReFold
base
ReFold
base
ReFold
base
ReFold
base
ReFold
Qwen3.6-27B (dense)
SWE-bench Verified
38
38
( ± 0)
50
41
( − 18%)
965
505
( − 48%)
1.71
0.83
( − 52%)
19.5
12.9
( − 34%)
SWE-bench Pro
5
6
( + 1)
46
50
( + 9%)
1097
899
( − 18%)
1.96
1.20
( − 39%)
20.8
20.5
( − 2%)
Multi-SWE-bench
25
25
( ± 0)
47
44
( − 6%)
1372
956
( − 30%)
1.55
0.93
( − 40%)
25.5
22.3
( − 13%)
Terminal-Bench 1.0
23
23
( ± 0)
20
15
( − 25%)
1279
596
( − 53%)
1.02
0.58
( − 43%)
27.4
19.0
( − 31%)
Table 1: Comparison across five benchmarks and two models, with 50 tasks per benchmark at the full context window and w=8 . Parentheses indicate changes relative to base .
Context, w
Arm
Resolve
Tokens (k)
Cost (¢)
Duration (min)
Compactions
KV peak (%)
Queue (s)
256 k, w=8
base
38
965
19.5
9.8
—
42
0
+ ReFold
38
( ± 0)
505
( − 48%)
12.9
( − 34%)
9.4
( − 4%)
—
29
( − 13)
0
( ± 0)
64 k, w=8
compact
37
882
18.5
9.8
0.1
43
0
+ ReFold
39
( + 2)
528
( − 40%)
14.1
( − 24%)
8.2
( − 17%)
0.1
( ± 0)
28
( − 15)
0
( ± 0)
32 k, w=8
compact
37
664
15.3
10.7
1.0
33
0
+ ReFold
35
( − 2)
374
( − 44%)
9.8
( − 36%)
9.3
( − 13%)
0.1
( − 92%)
30
( − 3)
0
( ± 0)
Table 2: Performance under context-budget and concurrency constraints. We vary the context budget at fixed w=8 , increase concurrency at the full context window, and combine both constraints. Each pair compares the reference baseline with and without ReFold; parentheses indicate the changes.
Figure 4: Ablation results. (a) Per-step context-management overhead on a log scale. (b) Peak KV-cache memory (left axis) and median duration (right axis) across chunk sizes on 20 tasks; the axes are scaled to coincide at k=3 . (c) Chunking (top) and reversibility (bottom) switched off and on.
Table 3: Serving, decoding and harness settings shared by every run.
Benchmark
Task
Full set
Our 50 tasks
SWE-bench Verified
GitHub issues in Python repositories, the human-validated subset of SWE-bench
500 tasks
8 repositories, all Python
SWE-bench Pro
Long-horizon issues in application and developer-tool repositories, reference patches of 107 lines over 4.1 files on average
731 tasks (public set)
all 11 repositories; Go 19 , Python 14 , TypeScript 12 , JavaScript 5
Multi-SWE-bench
GitHub issues in Java, TypeScript, JavaScript, Go, Rust, C and C++ repositories
1,632 tasks
19 repositories; Rust 13 , C++ 10 , JavaScript 9 , C 6 , Java 5 , TypeScript 4 , Go 3
Terminal-Bench 1.0
Tasks in a Linux container through the terminal: scientific workflows, networking, games, data analysis, security
80 tasks
50 of the 80
Terminal-Bench 2.1
Terminal-Bench 2.0, a harder and better verified set, with 28 of its tasks fixed
89 tasks
50 of the 89
Appendix
Table 4: The five benchmarks. Our 50 tasks cover every repository of the SWE-bench Pro public set and seven languages of Multi-SWE-bench, so the results are not tied to Python repositories.
Metric
Meaning
Formula
Per task, from the agent
Resolve
tasks the official evaluator scores as solved, out of N
∑i1[task i solved]
Exhausted
tasks with a request longer than the usable context B
∑i1[∃j:pij>B]
Timeouts
tasks ended by a model request that streams no token for 600 s
∑i1[task i timed out]
Steps
model calls of a task, median over tasks rounded to an integer
medini
Tokens
prompt and completion tokens of a task
N1∑i∑j(pij+gij)
Appendix
Table 5: Definitions of the metrics used in every table and figure.
Arm
Context
Run 1
Run 2
Lost
Gained
compact
16 k
38
35
5
2
recall
16 k
27
24
7
4
mask
16 k
17
25
5
13
compact
64 k
38
37
3
2
compact + ReFold
64 k
37
39
2
4
base
256 k
37
39
3
5
Appendix
Table 6: Seven configurations run twice on the same tasks. Lost and gained are the tasks solved only in the first or only in the second run, and six of the seven pairs differ by 1 to 3 tasks.
Strategy
Acts on
Removal decided by
Left in its place
As in
base
nothing
—
—
harness default
clip
observations older than the last three
a 1,000 -character cap
a truncation marker
harness default
mask
whole observation
recency: older than the last N=10
a one-line placeholder
observation masking ( Lindenbauer et al., 2025 )
prune
lines within observations
a trained 0.6 B scorer
the lines it keeps
SWE-Pruner ( Wang et al., 2026 )
recall
whole observations
age, once the context passes 75% of the budget
an addressable placeholder; recall <id> brings it back
ARC ( Dang et al., 2026 )
compact
the whole history, including agent messages
an extra model call, once the context passes 75% of the budget
a generated summary
Claude Code, Codex
Appendix
Table 7: The six strategies of Section 2 , by what they act on and what decides a removal. Only compact touches the agent’s own messages and keeps every task within the budget.
Arm
Setting
Resolve
Exhausted
Tokens (k)
KV cache (GiB)
Cache hit
Dur.
base
—
4
43
288
1.12
87.4%
5.2
clip
cap 4000
7
41
302
1.11
86.0%
6.0
cap 2000
6
42
325
1.12
82.4%
6.9
cap 1000
13
33
358
1.11
77.1%
8.0
mask
N=40
4
45
310
1.12
80.8%
6.1
N=20
13
31
410
1.09
30.0%
15.5
Appendix
Table 8: clip and mask at three settings of their knob at a 16 k budget. Each resolves the most at its tightest setting, still under half of what compact resolves.
Arm
Resolve
Steps
Tokens (k)
Completion (k)
Cost (¢)
Duration (min)
base
13
18
273
4.9
38.3
3.2
ReFold
12
( − 1)
15
( − 17%)
208
( − 24%)
4.5
( − 8%)
33.9
( − 11%)
2.6
( − 19%)
Appendix
Table 9: Claude Opus 4.8 on 21 SWE-bench Verified tasks at w=8 . ReFold cuts tokens, steps, cost and duration by 11 to 24% .
Figure 5: KV cache per session at every step with Qwen3.6-27B. The gap between base and ReFold widens with trajectory length, as base keeps everything it has read.
Figure 6: Tokens per task for every cell of Table 1 . ReFold shifts the whole distribution toward fewer tokens rather than trimming a few long tasks.
Cost (¢)
Method
Resolve
Tokens (k)
main
aux
total
Dur. (min)
KV cache (GiB)
base
38
965
19.5
—
19.5
9.8
1.71
LLMLingua-2
35
885
( − 8%)
18.6
0
18.6
( 0.95× )
10.7
( 1.09× )
1.52
SWE-Pruner
35
905
( − 6%)
18.7
0
18.7
( 0.96× )
9.6
( 0.98× )
1.55
AgentDiet θ=500
38
980
( + 2%)
20.6
4.5
25.1
( 1.29× )
20.8
( 2.12× )
1.72
AgentDiet θ=150
37
817
( − 15%)
17.5
12.2
29.7
( 1.52× )
34.0
( 3.46× )
1.45
Appendix
Table 10: Published methods at the setting of Table 1 . ReFold cuts tokens and cost the most, and only ReFold halves the KV cache of a session.
Dedup.
Fold
Resolve
KV cache (GiB)
Tokens (k)
—
—
38
1.71
965
✓
—
40
( + 2)
1.43
( − 16%)
649
( − 33%)
—
✓
36
( − 2)
0.88
( − 49%)
505
( − 48%)
✓
✓
38
( ± 0)
0.83
( − 52%)
478
( − 50%)
Appendix
Table 11: Deduplication and folding switched off in ReFold one at a time. The row with both off is base of Table 1 , and the row with both on is ReFold. Folding does most of the reduction, and the two together leave the smallest KV cache.
Routing, w
Arm
Resolve
Tokens (k)
Cache hit (%)
Duration (min)
Sticky, w=16
base
40
1118
89.5
12.5
+ ReFold
36
( − 4)
502
( − 55%)
76.4
( − 13.2)
9.5
( − 24%)
Round-robin, w=16
base
36
784
74.1
18.2
+ ReFold
36
( ± 0)
683
( − 13%)
63.9
( − 10.2)
15.6
( − 14%)
Sticky, w=32
base
38
898
78.3
23.1
+ ReFold
35
( − 3)
523
( − 42%)
72.2
( − 6.1)
15.3
( − 34%)
Appendix
Table 12: Qwen3.6-27B on two replicas under sticky and round-robin routing. Round-robin lowers the cache hit of both arms, and at w=32 , where base fills both KV pools, ReFold resolves 40 tasks against 26 .
Models, w
Arm
Resolve
Timeouts
Tokens (k)
Duration (min)
27B only, w=16
base
35
2
1033
27.0
+ ReFold
36
( + 1)
1
469
( − 55%)
15.4
( − 43%)
Alternating, w=16
base
33
3
1218
15.9
+ ReFold
37
( + 4)
0
499
( − 59%)
11.8
( − 26%)
27B only, w=32
base
21
21
667
61.6
+ ReFold
33
( + 12)
2
585
( − 12%)
37.9
( − 39%)
Appendix
Table 13: Qwen3.6-27B alone or alternating with Qwen3.6-35B-A3B at every call. ReFold keeps its token and duration savings under the alternation and resolves more tasks than base at both loads.
Memory
ReFold
Resolve
Exhausted
Tokens (k)
KV cache (GiB)
Steps
Duration (min)
—
—
4
43
288
1.12
31
5.2
✓
—
8
39
307
1.12
29
14.6
—
✓
21
14
334
0.87
32
6.9
✓
✓
17
23
616
0.98
52
21.8
Appendix
Table 14: Retrieval memory and ReFold switched independently at a 16 k budget. ReFold beneath the memory layer doubles its resolve, from 8 to 17 , and cuts its exhaustions from 39 to 23 .