AI Engineer Path

Week 28

Context engineering + Project 4

Decide exactly what occupies the context window, and why.

Why this week matters

The question has shifted from 'how do I phrase this' to 'what exactly is in the window right now, and why'. That's an engineering problem, not a writing problem, and it's where backend instincts apply.

Done when

Project 4 ships: document → validated Pydantic JSON, with retries, rate limiting, timeouts and cost logging. Build it the way you'd build a production service.

Milestone: Project 4: schema-validated extraction CLI

Concepts

5 lessons · tick each one once you could explain it

A context window holds: the system prompt, tool definitions, retrieved documents, conversation history, tool results, and room for the output. Make a budget: how many tokens each section may use, and what gets cut first when it overflows.

More context is not free accuracy. Irrelevant material dilutes attention, increases cost and latency, and can actively mislead (a retrieved but wrong document). The goal is the smallest set of high-signal tokens that lets the model do the task.

Measure it: log token counts per section per request. You can't manage what you don't measure.

text
Window: 200k tokens. Budget for one support-agent call:
  system prompt + rules ........  2k  (cached)
  tool definitions .............  3k  (cached)
  retrieved docs (top-5) .......  6k
  conversation history .........  8k  (compact beyond this)
  latest tool results ..........  4k  (truncate beyond this)
  reserved for output ..........  2k

Going deeper

Anthropic frames context as a finite 'attention budget' with diminishing returns: aim for the smallest set of high-signal tokens. Techniques include just-in-time retrieval (load data when needed via tools), structured note-taking outside the window, and sub-agents that return condensed results.

Long conversations and agent runs eventually exceed the window, and well before that they become expensive. Compaction summarises older turns into a compact state (decisions made, facts established, open tasks) while keeping the most recent turns verbatim.

Variants: rolling summaries; structured state objects (a JSON of what matters) instead of prose; external memory the agent can query on demand rather than carrying everything.

Risks: summaries drop details that later turn out to matter. Keep identifiers, numbers and commitments verbatim, and test compaction against conversations where an early detail is needed later.

Going deeper

Compaction quality is testable: plant facts early in a long conversation, compact, then ask about them. Track which categories of detail (numbers, names, decisions) survive.

Where this comes back

  • Week 33Long-term memory stores what compaction would otherwise discard.

A search tool returning 50 full web pages, or an API returning 10,000 lines of JSON, can blow the budget in one step. Truncate to the relevant fields, paginate, summarise large results, or return a handle the model can query further.

Shape context consistently: clearly tagged sections, the most important material at the start or end (models attend less reliably to the middle of very long contexts), and stable ordering so prompt caching works.

Tell the model what was truncated ('showing 20 of 340 results') so it can ask for more instead of assuming completeness.

Going deeper

Return identifiers and summaries by default, with a separate 'get details' tool for drilling in. It's the pagination pattern of APIs, applied to an LLM's context window.

Backend engineer tip: This is API design for a reader with a token budget: return projections, not whole entities.

Best resources for this lesson

Prompt caching stores the processed prefix of a prompt (system prompt, tool definitions, a long document) so later requests sharing that exact prefix skip its prefill, reducing input cost and time-to-first-token substantially. It requires the cached part to be byte-identical and at the start: put stable content first, volatile content (timestamps, user input) last.

Context rot: as context grows into the tens or hundreds of thousands of tokens, models get measurably worse at using it: missing details, confusing similar items, following stale instructions. A long window is a capacity, not a guarantee.

So: retrieve and compact rather than stuffing everything in, and test your feature at the context sizes it will actually see in production.

Going deeper

'Lost in the middle': models retrieve information placed at the start or end of a long context more reliably than in the middle. Chroma's context-rot study shows degradation grows with length even on simple tasks, especially with distractors.

Cache hits depend on exact prefixes and time-to-live windows; log cache-read vs cache-write tokens to verify your layout actually hits.

Common pitfalls

  • Putting a timestamp at the top of the system prompt, which breaks caching on every call.

Where this comes back

  • Week 26Caching skips the prefill phase for the shared prefix.

Build the extraction tool as a service would be built. Validation loop: Pydantic schema, structured output, retry with the error fed back, capped attempts. Rate limiting: a token bucket or semaphore that respects the provider's requests-per-minute and tokens-per-minute limits. Timeouts and backoff on every call, retrying only transient errors.

Observability: log for every call the model, prompt version, input and output tokens, cost, latency, retry count and validation outcome. Aggregate per document and per run.

Ship it as a CLI that processes a folder of documents concurrently and writes JSON plus a run report (success rate, total cost, p50/p95 latency).

python
@dataclass
class CallRecord:
    prompt_version: str; model: str
    input_tokens: int; output_tokens: int
    cost_usd: float; latency_s: float
    attempts: int; valid: bool

async def run(paths, concurrency=5):
    sem = asyncio.Semaphore(concurrency)
    async def one(p):
        async with sem:
            return await extract_with_retries(p)   # timeout + backoff + validation inside
    results = await asyncio.gather(*(one(p) for p in paths), return_exceptions=True)
    write_report(results)

Going deeper

Distinguish rate-limit dimensions: requests per minute, input tokens per minute and output tokens per minute. A token-aware limiter estimates a request's tokens before sending.

Fallback chains (primary model → secondary provider → cached or degraded response) plus timeouts per attempt keep features available during provider incidents.

Backend engineer tip: Everything you know about resilient service clients applies directly. Most people building with LLMs have never done this.

Where this comes back

  • Week 36These logs become your observability and cost dashboards.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Measure your context

    Instrument a chat app to log tokens per context section per request, and chart where the budget goes over a 20-turn conversation.

  2. Core

    Project 4

    Ship the extraction CLI: concurrency limits, token-aware rate limiting, timeouts, retries, validation retries, per-call cost logs and a run report.

  3. Stretch

    Context-rot test

    Plant a fact at different depths in 10k, 50k and 100k-token contexts with distractors, and measure retrieval accuracy for one model.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 190Monday29 Mar2 h planned

    Context window budgeting: what occupies it and why

  2. Day 191Tuesday30 Mar2 h planned

    Conversation compaction and summarisation

  3. Day 192Wednesday31 Mar2 h planned

    Tool-result truncation and structured context layouts

  4. Day 193Thursday1 Apr2 h planned

    Prompt caching and context rot at long contexts

  5. Day 194Friday2 Apr2 h planned

    PROJECT 4: document -> validated Pydantic JSON, retry on schema failure

  6. Day 195Saturday3 Apr3 h planned

    PROJECT 4: add rate limiting, timeouts, cost logging

  7. Day 196Sunday4 AprReview

    Review the week, finish anything unfinished, rest

Watch

Context Engineering for AgentsPrimary

LangChain

Project 4

Schema-validated extraction CLI

Document in, validated Pydantic JSON out, built the way you would build a production service.

  • Retries on schema failure with the validation error fed back
  • Rate limiting, timeouts, exponential backoff
  • Token and cost logging per call
  • Prompt library with versioned prompts, not one-off strings
Track it on the Projects page

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

What is context engineering, and how is it different from prompt engineering?
  • Deciding what occupies the window: instructions, tools, retrieved data, history
  • Budgeting, compaction, truncation, just-in-time retrieval
  • Engineering for cost, latency and quality, not just wording
How does prompt caching work and how do you design for it?
  • Reuses processed prefixes across requests
  • Stable content first, volatile content last
  • Monitor cache hit rates

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.For prompt caching to hit, the cached content must be…

  2. 2.Context rot refers to…

  3. 3.A tool returns 40k tokens of JSON. Best approach?

  4. 4.When compacting history, what should stay verbatim?

  5. 5.Which belongs in Project 4's per-call log?