AI Engineer Path

Concept Lab · Week 26 · AI Engineer Track

Prefill, decode and the KV cache

See why input and output tokens cost differently, and what caching saves.

The idea: Inference: prefill, decode and the KV cache

Prefill: all input tokens are processed in parallel in one forward pass, computing and caching each layer's keys and values. It is compute-bound and fast per token. It determines time to first token (TTFT).

Decode: each new token requires a full forward pass for just that position, attending to all cached keys and values. It is sequential and memory-bandwidth-bound. It determines tokens per second, and is why long outputs are slow and why output tokens are priced higher.

The KV cache stores keys and values for every previous token, so each decode step is incremental instead of recomputing the whole context. It grows linearly with context length, often dominating GPU memory at long contexts. Prompt caching reuses the prefill of a shared prefix across requests. Watch both phases in the simulation.

Open the full lesson in week 26

Next simulation: A RAG pipeline