Context Language Models
Organizations: University of Washington · Meta Superintelligence Labs · MIT · Trillium Labs
Abstract
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.
Figures & tables
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Input stream | Queries | Metric |
|---|---|---|---|
| Needle Retention | 4K-token chunks, each with 2–8 needle lines and 140 filler lines | none | needle lines retained verbatim in the final context |
| Sudoku Sketchpad | one move per turn on a board | current board after each move | board versions reproduced exactly |
| KV Store | batches of 100 SET operations with random 24-word values | 24 GET queries | exact-value accuracy |
| Log Triage | batches of 14–54 service log lines | 24 lookup and count queries | exact-answer accuracy |
| Method | Excerpt of the KV Store skill |
|---|---|
| CLM | Mechanism: Your context is mirrored to a file you can edit; changing that file changes what you are holding. Each batch: On the SAME turn you read a SET-BATCH, move the whole block out of your context and onto disk with this one command [a python3 heredoc that moves the block to /tmp/ctx_offload/ and leaves one placeholder line]. Each GET: Recover the value with grep -h ’SET <key> =’ /tmp/ctx_offload/* and then print the ANSWER block. Check: No note, or a token count that did not drop, means the batch is still in your context. |
| ACM | Mechanism: manage_context takes everything since your previous manage_context call … and replaces it with one message [summary_id: N] <summary>. Each batch: Read the gauge after every batch; when it is above 22,000, call manage_context on your very next turn, before you release another batch. Each GET: Spend one turn on query_memory(N, “the SET line for key K00084, verbatim”) with N the summary whose range covers the key. Check: What confirms a compression is the [summary_id: N] message and a [context: ~N/M tokens] readout lower than the turn before. |
| RLM | Mechanism: The REPL namespace persists: every variable and function you define survives from one operation to the next. That persistence is your store. Each batch: On the first operation define the store and the handler once; on every later operation only call them [a dictionary filled by a regular expression over the SET lines]. Each GET: A GET is answered by placing the ANSWER block in answer["content"] . Check: print(out) shows stored N keys so far growing by one batch per SET operation. |
| Context Folding | Mechanism: A branch begins from a copy of a context that already holds every batch and cannot delete anything from MAIN, so MAIN is as full after it as before. Each batch: On a SET-BATCH, run one command in MAIN and nothing else: echo READY_FOR_NEXT_OP. Do not open a branch. Each GET: Answer in MAIN … by finding the SET <key> = line in the batch in front of you and printing its value. Check: The [context: ~N/M tokens] readout should rise by the size of each batch and by almost nothing else. |
| Self-Compact | Mechanism: Compression happens only when C1=Y, C2=Y, C3=Y, N1=N. Then your whole history is replaced by … your summary. Each batch: On the SAME turn a SET-BATCH arrives, copy the whole block … to its own file. Answer [the probe] so that compression fires as soon as everything you have seen is on disk. Each GET: grep -h " ˆ SET K00084 = " /tmp/store/*.txt , then print the value in an ANSWER block. Check: What confirms a store is the count grep -c prints back, equal to 100. |
| Summary | Mechanism: At three quarters of the budget (about 24.6K of 32,768 tokens) … everything except the system prompt and this task message is then replaced by one message. Each batch: On the SAME turn a SET-BATCH arrives, copy the whole block … to its own file. Write the summary as a pointer, not an inventory. Each GET: grep -h " ˆ SET K00084 = " /tmp/store/*.txt , then print the value in an ANSWER block. Check: What confirms the store is the grep -c count coming back as 100. |
| Task | Category | Language | What the verifier scores | Start: files / LOC |
|---|---|---|---|---|
| ad_placement_optimization | Combinatorial optimization | C++ | solution score over judge cases | 1 / 20 |
| apple_incremental_game | Combinatorial optimization | Python | solution score over judge cases | 3 / 101 |
| graph_node_classification | Science & ML | Python | held-out accuracy (CPU-only judge) | 2 / 418 |
| grid_turing_robot | Combinatorial optimization | Python | solution score (lower is better) | 3 / 440 |
| juliet_vulnerability_analyzer | Software engineering | Python | hidden evaluator on the Juliet suite | 1 / 9 |
| openrct2_theme_park_ai | Games & simulators | JavaScript plugin | park value in a headless OpenRCT2 run | 6 / 5,950 |