On 29 September 2026, researchers from the University of Washington, Meta Superintelligence Labs, MIT and Trillium Labs posted "Context Language Models" on arXiv. The idea fits in one line: the model's context is saved as a file, and the model is allowed to edit that file however it likes. What if an agent could tidy its own working memory, instead of waiting for a fixed rule to do it?
In this blog, we will learn how Context Language Models work, and what the paper's results say about them. We will draw the hard parts as simple pictures, and keep a few of the paper's own figures where they explain the design best.
Paper at a Glance
Let me tabulate the basic facts of the paper for your better understanding.
| Item | Detail |
|---|---|
| Title | Context Language Models |
| Authors | Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Pang Wei Koh |
| Groups | University of Washington, Meta Superintelligence Labs, MIT, Trillium Labs |
| Posted | 29 September 2026 |
| Paper | arXiv:2609.37725 (CC BY 4.0) |
| Code | facebookresearch/context-language-models |
| Main models tested | Qwen3.6-27B, Qwen3.5-9B, Claude 4.6 Sonnet, GPT-5.4, GPT-5.6-Sol |
All numbers in this review come from the paper. We did not rerun the experiments.
What Is the Context of a Language Model?
The context is everything the model can see when it writes its next answer. In simple words, it is the model's working memory for one conversation.
For an AI agent, the context holds the system prompt, the task, every tool call and every tool result so far. Each turn adds more text to the end. A search agent, for example, adds ten search snippets one turn, a long document the next turn, and so on.
The problem is that the context has a size limit. Many of the paper's runs use a 32K-token limit. Once the context is full, the agent cannot read anything new. And long before that, old and useless text keeps costing compute on every turn.
Let's see this as below.

Here, we can see two agents doing the same three searches. On the left, a standard LM keeps every old result, so by turn 3 it passes the limit. On the right, the CLM replaces the results it has already used with a short note, so it stays under the limit. That right side is the whole idea of the paper.
How Do AI Agents Manage Their Context Today?
Agents today already shrink their context. The question is who decides when and how. The paper sorts the existing methods into a ladder of control, so let's see it as below.

Here, we can see three steps, with the model getting more control at each one.
The first step is harness rules. The harness is the program that runs the agent loop around the model. It decides when to shrink the context. A Codex-style summary harness, for example, summarizes the history when the context reaches 75% of the budget. MEM1 rewrites the context at every turn. The model writes the summary, but it never picks the moment or the method.
The second step is fixed tools. Here, the model decides when to act, but it can only pick from a small menu that people designed. Self-Compact lets the model decide when to compact. ACM adds tools to offload text to storage and fetch it back. Context Folding lets the model branch into a side task and fold it into a summary when it returns. The menu grows over the years, but it is still a menu.
The third step is the paper's own method, which removes the menu. We will open it in the next section.
There is also a related idea called Recursive Language Models (RLMs). An RLM keeps a long input in a REPL variable, so the model chooses what to read. But whatever it reads is still appended to the live context, so the context keeps growing. RLMs control what comes in. They do not clean up what is already there.
What Is a Context Language Model?
A Context Language Model (CLM) is a language model that manages its own context. In simple words, the model decides what stays in its working memory, what gets shortened and what gets thrown away.
The paper writes this as two short formulas. A standard LM only appends its new output to the old context:
Here, is the context at turn , and means "glue to the end". So the context can only grow.
A CLM instead builds the whole next context itself:
Here, can be any change the model wants. It can delete a block, rewrite a block in place, move a block to disk or add something new. The paper points out that this one formula covers every older method. A summary, an offload or a fold is just one particular choice of .
The authors link this to The Bitter Lesson, Rich Sutton's essay. That essay argues that general methods which let a system search and learn tend to beat rules written by people. Instead of hand-writing a context policy, the paper lets the model find its own.
Let's look at the paper's own overview figure as below.
Note
Figures labelled "Figure N from the paper" are reproduced from Shao et al. (2026), arXiv:2609.37725, under the CC BY 4.0 license. All other figures in this post are drawn by us.

Here, we can see the paper's whole story on one page. The top strip is the core idea: context goes in as a file, any function of it comes out as the next file . Panel (A) shows behaviors the model invented on its own. Panel (B) shows results without any training. Panels (C) and (D) show the two ways a CLM can learn better habits: in its prompt, and in its weights. We will walk through each panel in the sections below.
How Does "Context as a File" Work?
The paper builds CLMs with a simple trick. The model's live context is copied into a file, and the file path goes into the system prompt. The model can then edit that file with ordinary Bash commands, the same way it edits any other file.
What makes this file special is the sync. Whenever the file changes, the change is sent to the LLM server, and the next turn reads the edited file as its prompt. If the model does not edit the file, its new tokens are simply appended, like a normal LM.
Let's see the loop as below.

Here, we can see the three steps. First, the model writes a Bash command, for example a short Python script. Second, that command edits the context file: in the picture, two bulky tool results are swapped for one short "notes" turn. Third, the server reads the edited file as the next prompt.
Why a plain file and not a new set of tools? Because a file plus Bash can express any edit. The paper calls this its "less is more" design: keep as little as possible fixed in the harness, and let the model decide the rest. Even the search tools for deep research are given as in-context skills, not as a fixed tool list.
The same trick scales to many agents. An agent swarm starts with several context files, one per agent, each synced with its own server. A subagent is even simpler: the model starts one by creating a new context file, and stops it by deleting that file.
What Did CLMs Do With Their Own Context?
Once the model had full control, it came up with habits that no human had put on a menu. The paper collected these examples from its zero-shot runs, so no training was involved.
- Scoreboards for subagents: while running subagents on a math problem, the model kept a state block at the top of its context. The block held the best score, the number of agents and the key files. It updated this block through 163 in-place edits, while the whole context stayed at only 6-8K tokens.
- A new chat role: chat templates normally have roles like system, user and assistant. While rewriting its context, the model created a new role called "notes" for its own internal notes.
- Loops that prune results: the model wrote
forandwhileloops that walk over old turns and replace long or useless search results with one line, such as "No relevant results." - Reusable functions: the model defined its own helper function to compact old search turns, and then called it again and again.
- Good summaries: it squeezed 21K tokens into a summary that kept only the facts needed for the answer. In another run, it kept a ranked list of untried ideas for later.
Let's see the reusable-function example from the paper's Figure 3 (d), taken from a BrowseComp-Plus run at step 109. The ... parts are cut in the paper itself:
progress = """[Search progress: VERIFIED ... NEXT: ...]"""
s = re.sub(..., progress, s)
def compact_turns(text):
return re.sub(..., lambda m: m.group(0).split('\n')[0]
+ '\n[search results - see progress note]', text)
s = compact_turns(s)
Here, we can see the model doing two things. First, it writes a progress note and puts it into the context string s. Then it defines compact_turns, which keeps only the first line of each old search turn and points to the progress note for the rest. In that run, the model called compact_turns 37 times. In simple words, the model wrote its own context-cleaning tool and kept using it.
What Is ContextBench?
Before building CLMs, the authors asked a simple question: how well do today's methods manage context when nothing else is hard? Real tasks mix context skills with reasoning and knowledge, so a failure is hard to explain. So, they built ContextBench, a test that checks context management and nothing else.
Let's see its four tasks as below.

Here, we can see that each task tests one context skill. Needle Retention checks whether the agent keeps key lines word for word. Sudoku Sketchpad checks small edits in place on a 16 × 16 board. KV Store and Log Triage check whether the agent can move data out of its context and fetch it back when asked.
The input arrives one operation at a time, as new user messages, so no harness can quietly drop it on the way in. The tasks need no clever reasoning. An agent that could keep everything would answer every question. So the score measures only how well the agent manages its context.
The bottom row shows the setup. The limit is 32,768 tokens, and 2,048 of those are kept free for the reply. "Context pressure" is the total input divided by that limit. At 1× a keep-everything agent just fills its context, and the paper pushes up to 24×.
All seven methods ran on GPT-5.4. The result is clear in the paper's Figure 2. None of the existing methods is perfect, even on these simple tasks. CLM stays at or near full accuracy across the pressure levels.
The paper explains why each older method breaks:
- Summaries lose or invent details, which hurts Needle Retention and Sudoku Sketchpad.
- Methods without in-place editing must rewrite the whole Sudoku board for every single move.
- Standard coding tools can save data to disk, but they cannot remove it from the live context, so KV Store and Log Triage still fill up.
These failures are why the paper argues for letting the model control its context fully.
Why Does Editing the Context Cost Extra Compute?
There is a catch in editing the middle of the context. To see it, we first need the KV cache.
When a model reads a prompt, it computes a key and a value for every token in every attention layer. This first pass is called the prefill. The server saves these keys and values in the KV cache, so the next turn does not have to compute them again. This reuse is called prefix caching. In simple words, if the start of the new prompt matches the start of the old one, the server skips the work for that matching part.
But prefix caching only works up to the first changed token. Each token's key and value depend on every token before it, and on its position. So if the model edits the middle of its context, everything after the edit must be prefilled again, even text that did not change.
So, the paper counts cost with a metric it calls prefix-reuse FLOPs:
In simple words, we pay for every token from the first mismatch onward, plus every token the model writes. This is a fair way to measure a CLM, because it charges the model for the re-prefill that its own edits cause.
The paper gives a worked example on Qwen3.6-27B, with a 20,000-token prompt and a 500-token reply. Let's see it as below.

Here, we can see the same turn with the edit in three places. When the turn only adds new messages, 18,000 tokens come from the cache and the turn costs FLOPs. That saves 87% compared to having no cache at all. An edit in the middle leaves 10,000 reusable tokens and raises the cost to FLOPs. An edit at the very start is the same as having no cache, and costs FLOPs, which is 7.7 times the append-only turn.
The constants in the bottom panel come from the model's shape. Qwen3.6-27B has 64 layers, but only 16 of them use full attention. The other 48 are Gated DeltaNet layers, a kind of linear attention. That gives 48.70 × 10⁹ FLOPs per token, plus 3.93 × 10⁵ FLOPs for each query-key pair in the full-attention layers. We recomputed all three turn costs from the paper's formula, and they match the paper exactly.
So, every edit has a price. A CLM saves compute only if the context it removes is worth more than the re-prefill it causes. The good news, as we will see in the results, is that CLMs still come out cheaper overall.
What Is Suffix Cache Reuse?
What if we could avoid most of that re-prefill? So, here comes Suffix Cache Reuse (SCR) to the rescue. In simple words, SCR keeps the cache for every token that survived the edit, not just the tokens before it.
Let's see the paper's figure as below.

Here, we can see a context made of three blocks, A, B and C. The model replaces B with a new block B'. Standard serving reuses the cache for A, prefills B', and then has to prefill C again, because C sits after the first mismatch. SCR reuses the cache for A, prefills only B', and reuses the old cache for C as well.
There is a trade-off. The saved states for C were computed when B was still in front of it. So C's cache still carries a little memory of the old B. That makes SCR an approximation of a full re-prefill, not an exact copy. The paper notes this stale memory can even help sometimes, because it keeps some information from the past.
The paper also shows how SCR works on a hybrid model like Qwen3.6-27B, which mixes the two kinds of layers we met above. Let's see it as below.

Here, we can see the two cases side by side.
On the left are the full-attention layers. Each token stores its own key and value, so SCR can move the cached entries for C to their new place after B'. Moving them changes their positions, so SCR also re-rotates the RoPE position code on the keys. RoPE is the method that tells the model where each token sits in the sequence.
On the right are the linear-attention layers. These do not store one entry per token. They keep one fixed-size state that every token updates in turn. There is nothing per token to move, so SCR takes a snapshot of the state from before the edit and continues from it. In these layers, B' is not recomputed at all. It still gets computed in the 16 full-attention layers, and from there it can affect later layers.
One edit can leave many surviving pieces after the edit point, and each moved piece adds a bit of approximation. So SCR moves at most pieces per edit, the longest ones first, and prefills the rest. The paper tested values from 1 to 64 on 64 BrowseComp-Plus questions and saw no drop in accuracy, with the savings mostly flat after . So it uses everywhere.
The result: on BrowseComp-Plus with Qwen3.6-27B, SCR matches standard SGLang serving while using 65.0% of its prefix-reuse FLOPs. That is a 35% cut in server-side compute.
SCR also helps outside CLMs. Many reasoning models strip the old thinking text from earlier turns. That strip forces a re-prefill of everything after it, even when nobody edits the context. Of the 7.8% of prompt tokens that SCR reused beyond normal prefix-cache hits, 5.3 points came from this reasoning strip and only 2.5 points from the model's own edits. So any server that runs reasoning models could use SCR.
How Can a CLM Learn Better Context Habits?
Because context management is now something the model does, the model can learn to do it better. The paper shows three ways, from cheapest to most expensive.
Steering a CLM With One Sentence
The simplest way is to just tell the model what we want. The paper adds one sentence to the task prompt and checks whether the model's behavior follows it. It tests three behaviors on BrowseComp-Plus with Claude 4.6 Sonnet:
- Compact the context once it passes a size we pick (16K, 24K or 32K tokens).
- Compact only at the boundary between sub-questions.
- Back up the context file before every edit.
The model follows each instruction. In the threshold test, the median context size at the first compaction lands at 16.0K, 23.6K and 30.9K tokens for the 16K, 24K and 32K instructions, as the paper's Figure 1 (c) shows. In simple words, one sentence changes the policy, with no change to the harness or the weights.
Evolving a Skill Document
The next way is to write a better instruction automatically. Let be a skill document: a page of context-management advice placed in the prompt. The paper searches for the that gives the highest reward, while keeping the model and harness fixed.
Here, is the agent's run on task when it follows skill , and is the score of that run. In simple words, we try many versions of the advice and keep the one that scores best.
Let's see the loop as below.

Here, we can see five steps that repeat. The agent runs training tasks with the current skill. A proposer model reads those runs and writes full rewrites of the skill. Each rewrite is screened on a few training tasks, and the good ones are scored on a development split. The best skill becomes the starting point for the next round. Only after the loop ends does the final skill run once on a held-out test split.
The paper runs two versions on ContextBench with a 32K budget. In assisted evolution, Qwen3.6-27B is the agent and Claude Fable 5.1 writes the skills. In self-evolution, Opus 5 plays both roles.
The gains in assisted evolution are large. The agent starts with no context advice at all. On the development split, the evolved skill raises accuracy from 22.3% to 83.8% on KV Store and from 0.0% to 100.0% on Log Triage. It also lifts Sudoku Sketchpad from 45.3% to 65.8%, and Needle Retention from 97.6% to 100.0%. On the held-out KV Store test split, accuracy rises from 38.3% to 74.2%. That is the 35.9-point gain in the paper's abstract. Self-evolution with Opus 5 starts between 94% and 100% accuracy. Its evolved skills either cut cost at the same or higher accuracy, or raise accuracy further.
Reinforcement Learning With a Success-Gated Efficiency Advantage
The last way is to train the habits into the weights with reinforcement learning (RL). The paper uses GRPO, a popular RL method for LLMs. GRPO samples a group of answers for the same prompt and scores each one against the group average.
CLMs need one change. A CLM run is not one long, growing text, because the model keeps editing its past. So the paper uses stepwise GRPO. It cuts each run into its model calls, gives every call the score of the whole run, and trains on each call with the exact input it saw at that time.
The paper also wants runs that are both correct and cheap. Rewarding "fewer tokens" directly would be a trap. The model could learn to delete useful facts, or to make pointless edits that break the cache. So, the paper adds a success-gated efficiency advantage. In simple words, only the runs that solved the task compete on cost.
The rule is:
Here, is the prefix-reuse FLOPs of run , is the set of successful runs in group , and is their average cost. If fewer than two runs succeed, every efficiency advantage is 0. The final advantage is , where is the usual GRPO advantage from the task reward.
Let's see it with toy numbers as below.

Here, we can see four runs for one prompt. Runs 1, 2 and 3 solved the task, and their average cost is 1.50. Run 1 is cheaper than that average, so it gets a bonus of +0.20. Run 2 is more expensive, so it gets −0.20. Run 4 is the cheapest of all, but it failed, so it gets no efficiency bonus. The weight 0.25 is the value the paper used, so the cost signal only nudges the ranking among correct runs. It never makes a wrong answer look good.
The paper trains Qwen3.5-9B on 3,040 OpenResearcher deep-research prompts for 70 steps, with 8 prompts and 32 runs per prompt at each step. It uses 16 H200 GPUs for training and 48 for generating runs. It also trains a Codex-style summary harness with the same recipe, for a fair comparison. Then it tests both on BrowseComp-Plus.
Let me tabulate the results for your better understanding.
| Method | Accuracy before → after RL | PFLOPs per question before → after RL |
|---|---|---|
| Summary harness | 34.7% → 42.1% | 4.01 → 2.19 |
| CLM | 28.8% → 42.5% | 1.52 → 1.34 |
Here, we can see an interesting flip. Before training, the 9B CLM is about six points behind the summary harness, because a small model is weaker at managing its own context. After RL, the CLM gains 13.7 points to 42.5%, just above the trained summary harness. And it does this with 1.34 PFLOPs per question instead of 2.19, which is 38.8% less compute.
What Results Does the Paper Report?
All the main results below are zero-shot. The paper uses existing models as CLMs with no training at all. It runs every baseline out of the box too, so no method gets a training advantage.
Coding and Deep Research
The first test uses BrowseComp-Plus, a deep-research benchmark with 830 questions, plus two terminal-coding benchmarks: TerminalBench 2.1 and TBLite. Every method uses Qwen3.6-27B with a 32K context limit and a 100-turn cap.
The paper plots all seven methods in its Figure 5. To keep things simple, we redrew it with just CLM and the strongest baseline, the Codex-style summary, using the values from the paper's own figure. Let's see it as below.

Here, we can see two questions for each benchmark: how often the method is right, and how much compute it spends on each question. A good method has a longer bar on top and a shorter bar below. CLM wins or ties on accuracy everywhere, and it spends less compute everywhere. The other five baselines all scored below CLM on all three benchmarks.
- BrowseComp-Plus: CLM scores 59.4%, the best of all methods. That is 11.4% higher, in relative terms, than the strongest baseline, the Codex-style summary. It also uses 21.5% fewer FLOPs than the summary harness and 28.9% fewer than MEM1.
- TerminalBench 2.1: CLM matches the summary harness's accuracy while using only 70% of its FLOPs.
- TBLite: CLM beats the summary harness, 73.7% against 67.0%, with 91% of its FLOPs.
Model size matters here. With the smaller Qwen3.5-9B, CLM still beats the summary harness on BrowseComp-Plus, 39.9% against 37.7%. But the small model edits less. On TerminalBench 2.1, it edits its context 1.4 times per task and makes no edit in half the tasks, while Qwen3.6-27B edits 2.6 times per task. The 9B model's median peak context is 30.2K tokens of the 32K limit, against 17.6K for the 27B model.
Math Optimization
The next test uses four math problems from AlphaEvolve, where the agent writes a program and an evaluator scores it. The baseline is OpenEvolve, a workflow built only for this kind of evolutionary search. The CLM gets the same evolutionary method as plain advice in its prompt, and plans everything else itself. All methods use Claude 4.6 Sonnet with a 32K limit, and stop after 100 scored attempts or five hours.
Let me tabulate the best scores for your better understanding. OE stands for OpenEvolve and SA for subagents. Higher is better for the first three columns, and lower is better for Erdős overlap.
| Method | Circle packing ↑ | Heilbronn ↑ | Min-max/min-dist ↑ | Erdős overlap ↓ |
|---|---|---|---|---|
| OE | 2.541 | 0.03127 | 0.07690 | 0.38123 |
| OE-Agent | 2.525 | 0.03053 | 0.07724 | 0.38167 |
| CLM | 2.618 | 0.03653 | 0.07758 | 0.38094 |
| CLM (SA) | 2.636 | 0.03617 | 0.07758 | 0.38109 |
Here, we can see that a CLM, with or without subagents, has the best score on all four problems. The biggest gap is on the Heilbronn problem, where CLM beats OpenEvolve by 16.8%. A general agent with control over its own context beat a workflow built only for this job.
12-Hour and 24-Hour Repository Tasks
The longest tests run for hours. On EdgeBench-10, an agent improves one code repository for up to 12 hours, with a 32K budget and three seeds per task.
- With Qwen3.6-27B, CLM reaches a score of 44.6 using 179 prefix-reuse PFLOPs per trial. The summary harness reaches 42.3 using 437 PFLOPs. That is 5% higher with 59% fewer FLOPs.
- CLM with subagents reaches 44.2 at 181 PFLOPs, so subagents add little on a single repository.
- With Claude 4.6 Sonnet, CLM and its subagent version reach 51.0 and 50.4, against 42.3 for the summary harness.
With a bigger 128K budget, subagents start to pay off: CLM with subagents reaches 50.2, against 47.3 for CLM and 47.8 for the summary harness.
The last test is Software World. Six agents, one per repository, work in parallel for over 24 hours to make a group of linked Python packages faster. The score comes from 17 benchmarks in four downstream packages that the agents never see. All agents use GPT-5.6-Sol with a 272K context. At the same API spend, the CLM swarm gets 65% more downstream speedup than the summary swarm.
Our Review: Strengths and Open Questions
We like this paper for three reasons.
The first is how small the change is. There is no new architecture and no new tool menu. There is just one file, Bash, and a sync to the server. Because the formula covers every older method, any trick a person could add as a tool is also something the model can do on its own.
The second is the cost metric. Prefix-reuse FLOPs charge the CLM for the re-prefill its own edits cause. A paper that only counted context length would make CLMs look cheaper than they are. This one counts the real price and still comes out ahead.
The third is ContextBench. A test that checks context skills alone makes the failure of each older method easy to see and easy to explain.
We also see some open questions, and the paper names most of them itself.
- Small models need help. Zero-shot, the 9B CLM trailed the summary harness by about six points on BrowseComp-Plus, and only caught up after RL. The method works best when the model is strong enough to judge what to keep.
- Models are poor at counting their own context. The paper finds that current models guess their context length badly when it gets long. So, in its runs, the CLM gets a reminder 2,048 tokens before the budget runs out. A CLM is not yet fully on its own.
- SCR is an approximation. The reused states carry memory of text that was edited away. Its test used only 64 questions, so we would like to see it checked at a bigger scale.
- Safety is a real worry. A model that can write its own context can also store a bad instruction there, and that instruction would stay across turns. The paper points to earlier cases where a model put unauthorized instructions into its own summary. Editable context is a new place for prompt injections to hide.
The paper ends with a direction we find exciting: take the strategies that today live inside harnesses, and train them into CLMs. If that works, a harness becomes something like a skill a model learns once and keeps.
This is how Context Language Models work. We started with the context as a model's working memory and saw how today's harnesses and tools shrink it under fixed rules. Then we met the CLM, which saves its context as a file and edits it with Bash. We saw why edits in the middle cost a re-prefill, and how Suffix Cache Reuse wins most of it back. Finally, we saw a CLM learn better habits from one sentence, an evolved skill document and RL, and beat fixed methods at lower cost.