AI-Driven Developmentbeginner8 min

The Context Window

A model's only working memory — a fixed token budget everything has to fit inside, all at once.

Imagine trying to answer a complex question while only being allowed to look at one page of notes at a time. Whatever isn't on that page may as well not exist. A language model lives under exactly this constraint. As we saw with large language models, a model has no memory between calls — so the text you hand it in a single request is, quite literally, everything it can think about. The container that holds that text has a fixed size, and that container is the context window.

A bounded budget of tokens

The window is measured in tokens — the same small text chunks the model reads and writes. Every model has a maximum: it can take in only so many tokens per call, full stop. Think of it less like a notebook you can keep adding pages to and more like a single whiteboard of a fixed size. You can write whatever you like on it, but once it's full, something has to be erased before anything new fits. Window sizes vary a lot between models, from tens of thousands of tokens to a million or more, but every one of them has a ceiling, and that hard ceiling shapes almost everything about working with these models.

How it works: what shares the space

A lot has to fit on that one whiteboard at the same time. There's the system prompt — the standing instructions that set how the model behaves. There's the conversation history — every previous turn, re-sent so the model appears to remember. There are the tool definitions that describe what actions are available. And there's whatever you've pasted in — files, error logs, documentation. All of it competes for the same fixed budget.

Here's the part that surprises people: because the model has no memory between calls, the whole conversation is re-sent on every turn. A session doesn't fill the window once; it refills it on every call, a little fuller each time. Step through a long session below. The numbers are deliberately small (a 40K-token window) so you can do the arithmetic. When the next turn won't fit, predict what the model loses before you look, then flip where the rule lives.

Note

In our stack — when Claude Code works on your project, its context window is constantly being assembled: the instructions that guide it, the conversation so far, the definitions of the tools and any MCP connections it can use, and the slices of your codebase relevant to the task. The Claude model behind it has a generous window, but it is still finite — so Claude Code is deliberate about which files it pulls in rather than dumping the whole repository onto the whiteboard.

When the window fills up

Long sessions inevitably bump against the ceiling. When the total — instructions plus history plus tools plus files — would exceed the limit, something has to give, and there are two usual responses. The simplest is to drop the oldest turns, as in the scene above. The other is to summarize: compress earlier parts of the conversation into a short recap that keeps the gist and frees up tokens. Summaries are gentler, but they're still lossy. The recap of a two-hour session might keep "we're refactoring billing" and lose "and don't touch /legacy".

Either way, detail disappears, and it's usually the oldest detail: the instructions you gave at the very start, which are often the most important ones. This is why a model can seem to "forget" something you mentioned long ago in a marathon session. It didn't forget so much as the note got erased from the whiteboard to make room. Even before the hard limit, a single sentence buried under tens of thousands of tokens of later chat can carry less weight than you'd expect.

Watch out

Rules said once, early, are the first to go. A constraint you mention in your first message lives in the oldest turn, which is exactly what gets dropped or summarized away when the window fills. Put rules that must always hold into the pinned instructions (a system prompt or project instructions file), which are re-sent with every call. For a one-off rule, restate it close to the request it applies to.

Check yourself

Two hours into a session, an assistant starts editing a folder you told it to leave alone in your very first message. What's the most likely cause?

Why this is worth understanding

Once you picture the window as a finite, shared space, a lot of behavior stops being mysterious and becomes manageable. Putting the most important instructions where they won't get crowded out, keeping irrelevant files out of the budget, and starting fresh when a conversation has grown bloated are all ways of respecting the limit. It also explains cost and speed: everything in the window is sent again on every call, so a bloated session tends to be slower and more expensive as well as more forgetful. Doing this deliberately — choosing what deserves a spot on the whiteboard and what doesn't — is its own discipline, which the context engineering lesson explores in depth.

Check yourself

Your context window is nearly full. Which of these is NOT taking up any of it?

Key takeaways

  • The context window is the fixed maximum amount of text (measured in tokens) a model can consider in a single call.
  • It is the model's only working memory — it holds the system prompt, the conversation, tool definitions, and any pasted files, all competing for the same budget.
  • Because the model has no memory between calls, the whole conversation is re-sent every turn, so a long session refills the window a little fuller each time.
  • When the window fills up, the oldest content is dropped or summarized first — often taking your earliest instructions with it.
  • Rules that must always hold belong in pinned instructions (a system prompt or project instructions file); otherwise, restate them near the request they apply to.
  • Managing what goes into that limited space is a real skill — the subject of context engineering.

Keep going