A context window holds: the system prompt, tool definitions, retrieved documents, conversation history, tool results, and room for the output. Make a budget: how many tokens each section may use, and what gets cut first when it overflows.
More context is not free accuracy. Irrelevant material dilutes attention, increases cost and latency, and can actively mislead (a retrieved but wrong document). The goal is the smallest set of high-signal tokens that lets the model do the task.
Measure it: log token counts per section per request. You can't manage what you don't measure.
Window: 200k tokens. Budget for one support-agent call:
system prompt + rules ........ 2k (cached)
tool definitions ............. 3k (cached)
retrieved docs (top-5) ....... 6k
conversation history ......... 8k (compact beyond this)
latest tool results .......... 4k (truncate beyond this)
reserved for output .......... 2kGoing deeper
Anthropic frames context as a finite 'attention budget' with diminishing returns: aim for the smallest set of high-signal tokens. Techniques include just-in-time retrieval (load data when needed via tools), structured note-taking outside the window, and sub-agents that return condensed results.
Best resources for this lesson
- ArticleEffective context engineering for AI agents (Anthropic) · the definitive write-up
- VideoContext engineering for agents (LangChain, video)