FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents
Organizations: M365 Research, Microsoft
Abstract
LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.
Figures & tables
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Axis | ACON | PAACE | FOCUS (Ours) |
|---|---|---|---|
| Training regime | Offline guideline refinement via failure analysis; distilled into student models | Offline evolutionary prompt search; distilled into SLM using large corpus | Fully training-free; adapts online for every trajectory |
| Context operation | Generative rewriting / summarization of observations and history | Learned rewriting and summarization | Verbatim selection of thought–action–observation spans |
| Retention signal | Optimized natural-language guideline | Next- -task relevance from an externally supplied plan | Self-generated counterfactual utility via draft-model rollouts |
| External plan | Not required | Required (benchmark decomposition or separate planner) | Not required |
| Compressor | Distilled student | Distilled student | Any off-the-shelf small model (gpt-4.1-mini, Qwen3-8B/14B, Phi-4) |