Blog
Context windows, tokens and compaction explained
The context window holds everything a model has read in a session, counted in pieces called tokens, and both filling it up and emptying it back out follow specific, checkable mechanics, not vibes.
A token is the unit a language model actually reads: not a word or a character but a sub-word piece, averaging around four characters of English text in common tokenizers. The context window is every token in play for a request: the system prompt, tool definitions, every file read into the conversation, and the back-and-forth so far. Performance degrades well before the window is full, because models recall information at the start and end of a long context far more reliably than information buried in the middle, an effect called lost in the middle. Compaction summarizes older turns to make room, at the cost of detail a summary cannot keep.
What a token actually is
A token is the unit a model reads and generates, and it is neither a word nor a character. Tokenizers built with byte pair encoding break text into common sub-word pieces: short, frequent words stay as one token, while longer or rarer words split into pieces, so "encoding" might become encod and ing. OpenAI's own byte pair encoding tokenizer, tiktoken, states the rough yield directly in its documentation: each token corresponds to about 4 bytes of text on average, which for ordinary English text works out to close to 4 characters per token.
The exact count for a given sentence depends on which tokenizer a model uses, and different model families, sometimes even different versions of the same family, can tokenize identical text into different counts. Treat 4 characters per token as a rule of thumb for estimating, not an exact conversion, and use a model's own token-counting tool when the number actually matters.
What fills the context window
The context window is every token a request carries, not just the words a person typed. Anthropic's documentation states it plainly: everything in the request counts toward the window, including the system prompt, every message in the conversation, tool results, and the definitions of the tools made available to the model. A coding agent adds more again: the content of every file it reads, the output of every command it runs, and its own reasoning all become tokens sitting in that same window.
None of that is visible in a transcript the way a conversation is. A single file read can cost more tokens than an hour of typed messages, and a verbose command's output, a stack trace, a dependency list, can quietly make up most of what a model is carrying before it has been asked anything new.
Why a long session degrades before the window is full
A full context window is not the first place things go wrong. Researchers Nelson Liu and colleagues measured how language models actually use long contexts and published the result as "Lost in the Middle" in 2023: accuracy on a task was highest when the relevant fact sat at the very start or the very end of the context, and fell substantially when the same fact sat in the middle, even in models built for long contexts.
For an agent session this matters well before any limit is hit. The instruction given at the start of a long task and the file being edited right now are easy for a model to weigh; the decision made forty minutes ago, buried in the middle of everything that has happened since, is the part most likely to get underweighted when it matters.
What compaction does
Compaction is the mechanism that keeps a session going once the context window fills up: instead of ending the conversation, the system replaces older turns with a generated summary and continues from there. Claude Code's documentation describes it directly: when a long session compacts, the conversation history is summarized to fit the context window, and the same pass runs automatically as a session approaches its limit so a full window does not simply end the session.
Right after a compaction pass, what remains is a mix rather than a clean slate: a structured summary of everything that was said, plus certain content that gets reloaded fresh from disk rather than summarized, such as the system prompt and a handful of the most recently touched files.
What compaction loses
A summary is lossy by construction, and the loss is not evenly spread. Claude Code's own description of a compaction pass is a good example: the system prompt and the project's instruction files reload from disk, and a handful of recently touched files are re-read. Everything else, the ordinary back-and-forth of the conversation and whatever context a tool call added earlier, is reduced to whatever the generated summary chose to keep.
That is the practical lesson, not a criticism of the mechanism: compaction has to compress something to free space, and what it compresses is exactly the turn-by-turn record where decisions usually get made and never written down anywhere else. A decision that was only ever spoken inside the conversation survives only as whatever line the summary gave it, if it gave it one at all.
Why durable memory has to live outside the conversation
If a decision only exists inside a conversation, it is one compaction pass or one closed session away from being a line in a summary instead of the reasoning behind it. The fix is not a bigger context window, because performance degrades before a window fills and every window eventually does fill. The fix is to write the fact down somewhere a session reads fresh on arrival rather than reconstructs from memory.
A file such as CLAUDE.md, or the vendor-neutral AGENTS.md convention, is the common version of this: instructions meant to apply every session, loaded in full before an agent does anything. CLAUDE.md is too big. What do you take out? covers the failure mode on the other side of that same file: stuffing it with facts about the system rather than instructions on how to behave, until it is too large to maintain and spends context on sessions that have nothing to do with the fact sitting in it.
How this works in Forkbench
Forkbench is a desktop app for running coding agents, and each piece of work gets a Thread with its own notes, kept separate from the conversation happening inside it. A note written mid-session, a decision, a dead end, a convention, outlives the session that wrote it and the compaction pass that would otherwise have reduced it to a line in a summary, because it was never only inside the conversation to begin with.
Any agent working in that Thread reads the same notes on arrival, whatever vendor it comes from, so what a Claude Code session worked out can be read by a Codex session opened in the same Thread later. Agents can file notes themselves as they go, which is the part that compounds: the next session inherits the working-out instead of a summary of it. Does Claude Code have memory between sessions? covers the mechanism in more depth, including how a note differs from CLAUDE.md and what does not carry over when an agent is switched.
Stated plainly: a note is only useful if something gets written to it, so an agent that is never told to file what it finds will not do so on its own. And Forkbench does not read, trim or summarize CLAUDE.md or AGENTS.md on anyone's behalf; it offers a second place to put the material a compaction pass would otherwise erase, not an automatic fix for an instruction file nobody is maintaining.
Related: Does Claude Code have memory between sessions?, CLAUDE.md is too big. What do you take out?, How Forkbench works: Threads, notes and the board
Frequently asked
What is a token in AI?
A token is the unit a language model actually processes, a sub-word piece rather than a whole word or a single character. OpenAI's tiktoken library states a rule of thumb directly: each token works out to about 4 bytes of text on average, close to 4 characters for ordinary English, though the exact count depends on the tokenizer a given model uses. Use it to estimate, and a model's own counting tool when the exact number matters.
What counts toward a model's context window?
Everything in the request, not just what was typed. The system prompt, every prior message, tool definitions, and any file or command output read into the conversation all count as tokens in the same window. For a coding agent, file reads and command output are usually a larger share of that total than the conversation visible on screen.
Why does an AI agent get worse in a long session, even before it runs out of context?
Because models recall information at the start and end of a long context far more reliably than information in the middle, an effect researchers measured directly and published as "Lost in the Middle" in 2023. A decision made partway through a long session is the part most likely to get lost, well before the window itself fills up.
What does context compaction do, and what does it lose?
Compaction replaces older turns in a conversation with a generated summary once a session approaches its context limit, so the session keeps going instead of stopping. What survives is a mix: the system prompt and a handful of specific mechanisms reload automatically, but most of the turn-by-turn conversation is reduced to whatever the summary chose to keep, which is why a decision that was only ever spoken, never written down, can disappear.
Where should an AI agent's memory actually live?
Outside the conversation it was formed in. A conversation is one compaction pass or one closed session away from being a summary instead of the reasoning behind it, so anything a future session needs to know belongs in a file read fresh on arrival, such as CLAUDE.md or AGENTS.md for instructions, or a note for a finding or decision made along the way (see shared memory between sessions).
Keep reading
Sources
- tiktoken: OpenAI's byte pair encoding tokenizer (GitHub)
- Anthropic: Context windows
- Liu et al.: Lost in the Middle: How Language Models Use Long Contexts (arXiv:2307.03172)
- Claude Code: How Claude remembers your project (CLAUDE.md and auto memory)
- Claude Code: Explore the context window (what survives compaction)