A language model can only work with what is in front of it when it answers. Context engineering is the practice of deciding exactly what that is, so the model has what it needs and as little as possible of what it doesn't.
What the context window is
Every time you send a message to a large language model (LLM), the app sends the model a block of text. That block is the context, and the most it can hold is the context window. It is measured in tokens, small chunks of text that are often a whole word and sometimes part of one.
The context usually holds far more than the message you typed:
- Instructions: the system prompt that sets the model's role and rules.
- Chat history: earlier messages in the conversation, sent again on every turn.
- Files and documents: code, specs or pages the app pasted in.
- Tool definitions and results: what tools the model may call, and what came back when it did.
- Examples: sample inputs and outputs showing the style or format wanted.
The model itself is stateless between calls. It doesn't remember your last message; the app sends the history again each time. That is why the context window is best thought of as working memory, not long-term memory. What the model learned in training is its general knowledge. The context window is the desk: whatever is on it right now is all the model can look at for this answer. On most models the reply it writes counts towards the same limit, so a packed desk also leaves less room to answer.
Why a bigger window isn't a free win
A small window forces hard choices: a long spec and a long chat history may not both fit. So bigger windows sound like the answer, and they do help. They leave room for more rules, more examples, whole documents and the results of many tool calls.
But a model doesn't give equal, perfect attention to every token it is handed. As the context grows, answers tend to get worse, even when the useful information is all there. This is called context rot: performance drops as more material is packed into the window.
A few things drive it:
- Attention is spread thinner. Every extra token competes for the model's attention with the ones that matter.
- Position matters. Research on long inputs found that models are often best at using information near the start or the end, and worse at using details buried in the middle.
- Distractors mislead. Text that looks relevant but isn't, such as an old version of a file or a similar but different error, can pull the answer the wrong way.
- Cost and speed. More tokens mean more to process on every call, so a bloated context is slower and more expensive as well as less accurate.
What context rot looks like
You rarely get an error message. Instead the model:
- quietly stops following an instruction it was given at the start;
- misses a detail that is clearly in a document you provided;
- repeats a mistake from earlier in the chat, because the failed attempt is still in the history;
- answers with full confidence while losing track of the actual task.
That last one is the dangerous part. A model with a messy context doesn't say it is confused; it just gets sloppier.
From prompt engineering to context engineering
Prompt engineering is about wording: how to phrase one request so the model does what you want. It still matters, but it only covers the part you type.
Context engineering covers the whole setup. It asks, for each call: what should the model see before it answers? That includes the instructions, which parts of the history to keep, which files to include, how much of each tool result to pass back, whether a summary would do, and which examples to show. The aim is the smallest set of tokens that still gives the model everything it needs.
The difference matters most for agents: a model that works in a loop, calling tools, reading their output and deciding what to do next. Each step adds more to the context, so without care an agent fills its own window with logs, file dumps and dead ends.
What each part of the context needs
Instructions
Write rules that are clear and specific but not a wall of edge cases. A long list of 'if this, then that' rules is hard to follow and easy to half-forget. Group related rules, and put the ones that matter most where the model is likely to see them.
History
Keep the turns that still matter, such as the goal, decisions made and constraints agreed. Summarise older turns into a few lines rather than resending every message. Drop failed attempts once you have learned from them, or they may be copied.
Files and documents
Include the relevant part, not everything that might be relevant. For a bug in one function, that function and its callers beat the whole repository. Many tools now fetch just in time: they keep a list of file paths or links, and load a file only when the model asks for it.
Tool results
A tool call can return far more than the model needs: a full web page, thousands of log lines, a large JSON response. Trim it to the fields that answer the question before it goes back into the context.
Examples
A few well-chosen examples usually teach format and style better than a paragraph describing them. Pick examples that show the typical case clearly, rather than piling in every unusual one.
A worked example: a bug-fixing assistant
Say you are building an assistant that fixes bugs in a codebase. A user has been chatting with it for an hour and now reports a new failing test.
The naive approach sends everything: the full chat, every file opened so far, and the complete output of the test run. It fits in a large window, but the fix for the new bug is buried among old attempts and unrelated files.
A context-engineered version builds the context for this request on purpose:
def build_context(task, history, repo, budget):
parts = [SYSTEM_RULES]
# Old turns become a short summary.
parts.append(summarise(history[:-6]))
parts.extend(history[-6:])
# Only the files that match this task.
for f in repo.search(task, limit=3):
parts.append(f.relevant_lines())
# The failing test, not the whole run.
parts.append(failing_tests_only(task.output))
# The task goes last, where it stands out.
parts.append(task.description)
return trim_to_budget(parts, budget)The helpers are made up, but the choices are the real point:
- The rules stay at the top, and the current task sits at the end, away from the middle where details are most easily missed.
- Old history is compressed, while the last few turns stay word for word.
- Three relevant files beat thirty possibly relevant ones.
- The test output is cut to the failure that matters.
- There is a token budget well under the model's limit, so the context never simply grows until it is full.
The model is the same in both versions. Only what it sees has changed, and that is usually what decides whether the fix is right.
Techniques for long tasks
Some jobs run for many steps, such as an agent working through a large refactor. Even careful choices eventually fill the window, so long-running systems use a few extra techniques:
- Compaction. When the context gets near its budget, summarise it: keep the goal, decisions, open problems and key file names, and start a fresh context from that summary.
- Structured notes. The agent writes progress to a file outside the context, such as a to-do list or a notes file, and reads it back when needed instead of carrying everything in the window.
- Sub-agents. A focused helper gets a clean context for one sub-task, such as searching the codebase, and hands back only a short result. The main agent's context stays tidy.
Each one trades a little detail for a lot of clarity. The risk is summarising away something that turns out to matter, so good summaries keep decisions and constraints, not just a vague 'progress so far'.
Common mistakes
- Treating the window as a target. Having room for a million tokens doesn't mean you should use it. Fill it only with what helps.
- Pasting whole files or whole logs when a few lines would do.
- Letting history grow forever, so early instructions drift out of the model's focus.
- Keeping failed attempts in the context, which invites the model to repeat them.
- Fixing everything with the prompt. If answers get worse as a session goes on, the wording is rarely the problem; the context is.
Key takeaways
- The context window is everything the model can see for one answer: its working memory, not its long-term memory.
- Bigger windows help, but performance tends to drop as they fill up. That is context rot.
- Context engineering means choosing what the model sees before it answers, across instructions, history, files, tool results, summaries and examples.
- Keep what is relevant, summarise what is old, and cut the rest.
- A messy context gives messy answers, whatever model you use.