An AI agent starts a task with clear instructions, a few relevant files, and one concrete goal.
Then it searches documentation. A tool returns a long page of markup. The agent runs a command and receives hundreds of log lines. Someone adds a correction halfway through the conversation. A retrieval system supplies five related documents, including an older version of the policy. The agent summarizes its progress, then that summary is appended beside the original transcript.
Nothing has exceeded the model's context window. Every piece of information fits.
The agent still gets worse.
It follows the stale instruction instead of the correction. It cites the superseded policy. It misses the one error buried in the logs. It spends time reconciling two summaries that describe the same decision differently.
The failure is not that the model ran out of room. The failure is that the room became difficult to use.
Capacity and Quality Are Different Problems
A context window is the amount of input a model can consider during one inference: instructions, conversation history, retrieved documents, tool results, and anything else the surrounding system sends.
A larger window raises the maximum capacity. That is useful. An agent can inspect a longer document, retain more of a conversation, or work across more tool calls before older material must be removed or compressed.
Capacity does not guarantee that every included fact receives equal attention or that the model will resolve contradictions correctly.
The earlier beginner explainer, Why AI Forgets What You Just Told It, describes what happens when working context is finite. Agent builders face a second problem before the limit is reached: deciding what deserves to occupy that working context at all.
Anthropic calls this context engineering: curating the set of tokens most likely to produce the desired behavior. Its current engineering guidance treats context as a finite resource with diminishing returns and recommends techniques such as just-in-time retrieval, compaction, structured note-taking, and clearing unnecessary tool results.
The goal is not to make context small for its own sake. It is to make each part useful, current, and legible to the model.
Five Ways Context Becomes Noise
1. Stale instructions remain active
An agent may begin with “do not change the database.” Later, after a review, it receives permission to update one specific record. If both instructions remain in a long transcript without a clear statement of current authority, the agent must infer which one governs the next step.
Sometimes the reverse happens: an early permission remains visible after it should have expired.
Instructions need scope, ownership, and a lifecycle. A current task brief should say what is allowed now. Superseded instructions should be removed, marked obsolete, or summarized into the current state instead of remaining as equally plausible text.
2. Retrieval returns related material, not authoritative material
Semantic search is good at finding text that resembles a query. Similarity is not the same as authority.
A search for a deployment procedure may retrieve a current runbook, an old incident note, a proposal that was never adopted, and a copied excerpt with no date. Giving the agent all four does not necessarily make the answer safer. It creates a source-selection problem inside the context window.
Retrieval should consider freshness, document status, source ownership, and task relevance—not only semantic similarity. When sources disagree, the context should make the conflict explicit rather than asking the model to notice it by accident.
3. Tool output arrives at the wrong resolution
Tools often return what is easy for software to emit rather than what the next reasoning step needs.
A browser tool returns an entire page when the agent needs one table. A shell tool returns ten thousand log lines when the agent needs the first failure and its surrounding events. A database tool returns every column when the task needs an identifier, state, and timestamp.
Raw output consumes context, hides signal, and may introduce unrelated instructions or secrets from the retrieved material. A well-designed tool should return bounded, structured evidence and a reference that lets the agent fetch more if needed. That is part of designing tools a model can actually use, not merely an optimization after the tool works.
4. Conversation history becomes duplicated state
Long-running agents often carry the original messages, later corrections, periodic summaries, and notes generated from those same messages.
Duplication is not neutral. Two summaries may compress the same decision differently. A fact repeated several times can look more important than a newer fact stated once. An old plan can continue to compete with the plan currently being executed.
The agent needs a canonical task state: goal, constraints, completed work, open decisions, and the next action. The transcript can remain available for audit, but it does not all need to stay in the model's immediate working set.
5. Context from one task leaks into another
An agent that moves between customers, repositories, or operational incidents can inherit assumptions from the previous task.
Even harmless residue creates risk: the wrong coding convention, environment name, document version, or approval state. Sensitive residue is worse.
Task-local context should be the default. Shared memory belongs in an explicit, reviewed layer. Starting a new task should not mean carrying the entire prior conversation forward because storage is convenient.

Long Context Can Hide Important Information
This is not only an intuition about clutter.
The research paper Lost in the Middle tested multi-document question answering and key-value retrieval. It found that model performance could change significantly depending on where relevant information appeared, with information in the middle of long inputs often used less reliably than information near the beginning or end.
That 2023 study does not describe every current model or every task. Models have improved, architectures differ, and providers continue to expand long-context capabilities. Its durable lesson is narrower: accepting a long input is not proof that a model will use every part of it robustly.
Agent systems should test their own context shapes. Put the same evidence at different positions. Add realistic distractors. Duplicate an older value. Include a large tool result before the key instruction. Then measure whether task success changes.
The question is not “Does the model support this many tokens?”
It is “Does the agent still make the right decision when the context looks like production?”
Build a Context Hierarchy
A useful context has visible structure and authority.
One practical ordering is:
- Current instructions and boundaries. What is the goal? What may the agent do? What requires approval?
- Task state. What is already complete? What remains open? Which decisions are current?
- Relevant evidence. Which sources directly support this step, and which source wins if they conflict?
- Bounded tool results. What happened, what identifiers changed, and where can details be retrieved?
- Recent interaction. Which turns are necessary to interpret the current request?
- References. Which files, traces, or records can be fetched if the next step needs them?
This is not a universal prompt template. It is a design principle: stable authority should be easy to distinguish from transient evidence and historical detail.
Ordering helps, but selection matters more. A perfectly labeled pile of irrelevant information is still a pile.
Retrieve Just in Time
The instinct to preload everything usually comes from fear that the agent will miss something later.
A better pattern is progressive disclosure. Give the agent a map—file names, source descriptions, timestamps, identifiers, indexes—and let it retrieve the specific content needed for the current decision.
OpenAI's agent-building guidance similarly describes retrieving the information needed for a request and reranking results before adding the most relevant chunks to context.
Just-in-time retrieval has its own failure modes. The agent may search poorly or fail to fetch a necessary source. That is why references need clear names and descriptions, retrieval needs evaluation, and critical instructions should not be hidden behind optional discovery.
The useful split is:
- preload the small amount of information that must always govern behavior;
- retrieve evidence that depends on the task and current step;
- keep bulky details outside the window until they are requested.
Compact Without Erasing Decisions
Compaction replaces a growing history with a smaller representation of what still matters.
A good compaction preserves:
- the active goal and constraints;
- decisions and who authorized them;
- exact identifiers and changed state;
- unresolved questions and failed attempts;
- evidence references needed for verification;
- the next intended action.
It can discard repeated status messages, obsolete plans, successful low-value tool output, and conversational phrasing that no longer affects the task.
Compaction is lossy. A summary can omit the detail that later becomes important. Keep the original trace outside the immediate context so a person or agent can retrieve it. The flight recorder is the audit record; the compacted task state is the working copy.
Test Context Changes Like System Changes
Context assembly is part of the product, not an invisible preprocessing step.
When you change retrieval count, summarization, tool-result formatting, memory rules, or instruction ordering, rerun representative tasks. Track outcomes, not just token use.
Useful tests include:
- Ablation: remove one context source. Does accuracy improve or fall?
- Conflict: include an old and new value. Does the agent choose the authoritative one?
- Placement: move the decisive fact. Does behavior change?
- Noise: add realistic irrelevant results. Can the agent ignore them?
- Growth: extend the task across many tool calls. When does performance begin to drift?
- Recovery: after compaction, can the agent resume without repeating or undoing work?
Capture which sources were selected, their versions, and how much content each contributed. That evidence belongs beside ordinary agent traces because it explains why the model saw what it saw.
Better Context Is Maintained Context
There is no permanent perfect context bundle.
Instructions change. Documents age. Tools evolve. Tasks move from discovery to execution to review. The information useful at one stage becomes noise at another.
The system has to maintain context over time: replace stale state, preserve authority, retrieve current evidence, bound tool output, compact carefully, and keep details recoverable outside the working window.
A larger context window is valuable capacity.
Reliability comes from deciding what deserves to enter it.