KV-streams for Efficient Compaction in Agentic Reinforcement Learning
Authors: Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, +10 more
Organizations: Mila · Microsoft · Université de Montréal · McGill University · Polytechnique Montréal · ServiceNow Inc · Rensselaer Polytechnic Institute · University College London, University of London · Vmax · Cohere · Edinburgh University · HEC Montréal
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.
Figures & tables
Figure 1: Final evaluation score against total training compute for every run on TextWorld, ALFWorld and SWE-bench Verified. Axis ticks give GPU hours and wall-clock hours. Colour denotes the compaction strategy, circles are KV-streams and squares re-prefill compaction, with full context as the reference. Error bars give the standard deviation over three seeds on SWE-bench Verified, and the standard error over seeds, or over evaluation episodes for single-seed runs, on the text-based games. KV-streams reaches the same final score as re-prefill compaction at a fraction of the compute.
Figure 2: Compaction strategies and their training costs. (A) re-prefill compaction recomputes retained tokens and creates separate training traces. KV-streams preserves retained KVs and trains on a continuous trace with an attention mask that reproduces eviction. (B) context transformations under each strategy (P: prompt; S: summary).
Figure 3: Extra tokens prefilled over a 32 k-token rollout vs. tokens kept per compaction o+s (overlap o plus summary s ), for memory budgets B . As the retained tokens approach the budget, the repeated prefill grows by a factor of B−(o+s)N , the number of times each kept token is re-processed by the trainer. KV-streams prefills nothing again (green, at 0 ).
Figure 4: KV-streams reaches comparable or better success with fewer GPU hours on TextWorld (left) and ALFWorld (right). Axis ticks give GPU hours and wall-clock hours ( 8 GPUs on TextWorld, 4 on ALFWorld). Solid and dashed curves denote re-prefill compaction and KV-streams, with full context as a reference. Arrowheads mark truncation after exceeding the runtime of full-context completion. We provide complete curves appear in Figure 10 .
Figure 5: Smoothed per-step timing and rollout statistics on TextWorld for runs that do not exceed the full-context wall-time. KV-streams (dashed) reduces generation and forward/backward time relative to re-prefill compaction (solid). Rollout statistics show mean turns per episode and sequence length. We find KV-streams to be substantially faster during forward/backward and generation.
Figure 6: Comparison of KV-streams against full context and re-prefill Markovian Thinker on software-engineering tasks. Left: SWE-bench Verified success rate, smoothed with a centred three-checkpoint average. The final point is the mean and standard deviation over three evaluation seeds of the last checkpoint, as in Figure 1 . Right: training reward, smoothed with an exponential moving average over 8 gradient steps for every run (raw trace faint). Axis ticks give GPU hours and wall-clock hours on 8 GPUs, and every run is shown up to 200 gradient steps. KV-streams reaches peak performance about 3× faster than full context and about 2× faster than re-prefill compaction.
Figure 7: (A) Transfer of TextWorld-trained checkpoints to other environments. Bars show mean reward and whiskers show the standard deviation across evaluation seeds. (B) An SFT warm-start does not pay for itself once its cost is counted. TextWorld success rate against GPU hours, as in Figure 4 , for Sliding-Window and Markovian Thinker with KV-streams, with an SFT warm-start followed by RL (SFT + RL, solid) and RL only (dashed). SFT + RL curves start after the GPU hours spent on SFT (SFT-Time arrow). The SFT loss is in Appendix C.2 .
Figure 8: Post-eviction retrieval in-distribution (A), on new names (B), and with TV trained but unlisted (C). Heatmaps show RL and SFT accuracy (mean ± SD over three seeds) at the best learning rate per condition. Panel C covers all trained targets; its inset compares empirical TV-answer probabilities in the initial and final training batches of an RL run. Icons represent text targets.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
TextWorld
ALFWorld
SWE
Model
Qwen3-4B-Instruct-2507
Qwen3-4B-Instruct-2507
Qwen3.5-4B
GPUs (inference / trainer)
4 / 4 H100
1 / 3 A100
4 / 4 GB200
Rollouts per prompt
8
8
8
Batch size (rollouts per step)
512
128
256
Gradient steps
500
200
200
Learning rate
10−6
10−6
10−6
Appendix
Table 1: Hyperparameters shared by all runs within each benchmark. Compaction strategy and rollout concurrency vary between runs (Table 2 ).
TextWorld
ALFWorld
SWE
Summary
48 ( 192 )
64 ( 64 )
–
Sliding-Window
48 ( 192 )
64 ( 64 )
70 ( 280 )
Markovian Thinker
48 ( 192 )
64 ( 64 )
68 ( 272 )
Markovian Picker
48 ( 192 )
64 ( 64 )
–
Full context
12 ( 48 )
16 ( 16 )
28 ( 112 )
Appendix
Table 2: Concurrent rollouts per inference GPU for each strategy, with the total across GPUs in parentheses. Settings are the same for re-prefill compaction and KV-streams.
Figure 9: Sampling throughput against the number of concurrent rollouts per GPU on TextWorld, measured once 512 rollouts have completed. Left: rollouts per second. Right: generated tokens per second. KV-streams uses a sliding window of 10 turns. Hollow markers denote settings where the KV cache fills and vLLM preempts requests.
Figure 10: Training curves through each run’s last checkpoint (dot). Summary and Sliding-Window with re-prefill compaction exceed the full-context compute budget. The lower rows show smoothed mean turns per episode and rollout length, excluding padding.
Figure 11: Per-step timing and rollout statistics, including Summary and Sliding-Window with re-prefill compaction.
Figure 12: Training curves by gradient step on TextWorld (left) and ALFWorld (right). Top: smoothed mean training reward. Bottom: evaluation success rate, with faint points for individual evaluations. Multi-seed curves show the per-step mean and shaded seed range. Solid: re-prefill compaction; dashed: KV-streams.
Figure 13: SFT warm-start loss by optimizer step. The faint line shows per-step loss; the bold line shows a centered 9 -step rolling mean.
Figure 14: Mean post-eviction accuracy over three seeds: (A) in-distribution, (B) prompt generalization, and (C) a zero-support target (TV trained but unlisted). Panel C averages all six targets. Smaller cell labels give KV-cache token capacity k .
Figure 15: Generalization to new actor names with (A) listed options and (B) no listed options. Cells show mean retrieval accuracy over three seeds; smaller labels give KV-cache token capacity k .
LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV cache compaction can reduce this cost, but most prior methods assume a static context where future queries are known or can be approximated offline. Agents instead require online compaction: new information must be compressed before future relevance is known, using proxy queries cheap enough for the inference path. We study online compaction across token eviction (TE) and attention matching (AM), adapting both to compact agent turns and comparing cheap proxy sources such as boundary, repeat-prefill, and delayed future-generation queries. Experiments on BrowseComp-Plus and WideSearch show that immediate compaction often hurts performance, whereas delaying compaction to use the agent's future queries recovers much of the gap. Moreover, TE is often more robust than AM under imperfect proxies. Across models at different scales, TE preserves most of the accuracy while reducing KV cache by 80%, and can improve throughput over the no compaction baseline. These results position proxy-query selection as a core design choice for practical online KV compaction.
Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizing previous interaction states and continuing the rollout under a compressed context, but incorporating compaction into reinforcement learning remains underexplored. We propose CompactionRL, a reinforcement learning strategy to train long-horizon agentic LLMs with context compaction. Our approach jointly optimizes task execution and summary generation with token-level loss normalization and cross-segment generalized advantage estimation. This design enables the LLM agents to learn from compacted long-horizon trajectories. We train CompactionRL on top of open models and observe consistent performance gains on agentic coding tasks. CompactionRL enables the open GLM-4.5-Air model (106B-A12B) to achieve Pass@1 scores of 66.4% on SWE-bench Verified and 26.2% on Terminal-Bench 2.0, exceeding the base model under inference-time compaction by 6.6 and 4.9 points, respectively. Built upon GLM-4.7-Flash (30B-A3B), CompactionRL improves Pass@1 by 5.5 and 6.7 points against the base model, reaching 56.0% on SWE-bench Verified and 20.2% on Terminal-Bench 2.0. CompactionRL is thus deployed in the RL pipeline for training the open GLM-5.2 model (750B-A40B).
Yujiang Li, Zhenyu Hou, Yi Jing +2
Work done while YL, ZH, and YJ interned at Z.AI. · Tsinghua University
We present PolyKV, a system in which multiple concurrent inference agents share a single, asymmetrically compressed KV cache pool. Rather than allocating a separate KV cache per agent -- the standard paradigm -- PolyKV writes a compressed cache once and injects it into N independent agent contexts via HuggingFace DynamicCache objects. Compression is asymmetric: Keys are quantized at int8 (q8_0) to preserve softmax stability, while Values are compressed using TurboQuant MSE -- a Fast Walsh-Hadamard Transform (FWHT) rotation followed by 3-bit Lloyd-Max quantization with centroids tuned to N(0,1). We evaluate across two model scales (SmolLM2-1.7B-Instruct and Llama-3-8B-Instruct), three context lengths (600-7,194 tokens), and up to 15 concurrent agents. PolyKV achieves a stable 2.91x compression ratio across all configurations. On Llama-3-8B with 15 agents sharing a 4K-token context, PolyKV reduces KV cache memory from 19.8 GB to 0.45 GB -- a 97.7% reduction -- while maintaining only +0.57% perplexity degradation and a mean BERTScore F1 of 0.928. PPL delta does not grow with agent count and improves as context length increases, inverting to -0.26% at 1,851 coherent tokens. To our knowledge, no prior work combines a single shared, lossy-compressed KV pool with multi-reader concurrent agent access.