cs.CLOct 6, 2026

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

Authors: Yupeng Su, Jiayi Tian, Zheng Zhang, Souvik Kundu

Organizations: University of California, Santa Barbara · Intel

Abstract

Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Context Recycling for Long-Horizon LLM Inference

    May 1, 2026Derek ThomasLLM Inference OptimizationDialogue Benchmarks

  2. FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

    Sep 29, 2026Shantanu Dixit, Anson Bastos, Xuchao Zhang +2Large Language Model AgentsInteraction History