Concept Lab · Week 26 · AI Engineer Track
Prefill, decode and the KV cache
See why input and output tokens cost differently, and what caching saves.
The idea: Inference: prefill, decode and the KV cache
Prefill: all input tokens are processed in parallel in one forward pass, computing and caching each layer's keys and values. It is compute-bound and fast per token. It determines time to first token (TTFT).
Decode: each new token requires a full forward pass for just that position, attending to all cached keys and values. It is sequential and memory-bandwidth-bound. It determines tokens per second, and is why long outputs are slow and why output tokens are priced higher.
The KV cache stores keys and values for every previous token, so each decode step is incremental instead of recomputing the whole context. It grows linearly with context length, often dominating GPU memory at long contexts. Prompt caching reuses the prefill of a shared prefix across requests. Watch both phases in the simulation.
Next simulation: A RAG pipeline