AI Engineer Path

Week 26

LLM 101

Speak the vocabulary of the role: training stages, tokens and cost, decoding, inference, memory.

Why this week matters

This is the vocabulary of every conversation you'll have as an AI engineer. The specialisation starts here, and it is 30% of the whole plan.

Hint for the done-when test: per token, output is priced higher, yet in real systems the prompt usually dominates. It is far longer, it is re-sent on every turn of a conversation or agent loop, and every extra context token makes each decode step attend over more keys and values.

Done when

You can explain why a long prompt costs more than a long completion.

Concepts

9 lessons · tick each one once you could explain it

Pretraining: next-token prediction on trillions of tokens. Produces broad knowledge and skills, but a base model that continues text rather than following instructions. It is by far the most expensive stage.

Supervised fine-tuning (SFT): train on thousands to millions of curated prompt→response conversations. The model learns the assistant persona and format. RLHF trains a reward model on human preference comparisons and optimises the LLM against it (classically with PPO). DPO skips the reward model and optimises directly on preference pairs: simpler and now common.

Newer stages use reinforcement learning with verifiable rewards (unit tests passing, maths answers matching) to train long reasoning. As an engineer you mostly consume these models, but knowing the stages explains behaviour: sycophancy, refusals and verbosity are post-training artefacts; knowledge gaps are pretraining artefacts.

Going deeper

RLHF in three steps: collect human comparisons of model outputs, train a reward model to predict preferences, then optimise the LLM against the reward with a KL penalty that keeps it close to the SFT model. DPO shows the same objective can be optimised directly on preference pairs without a separate reward model.

Constitutional AI replaces many human labels with AI feedback guided by written principles, the basis of RLAIF.

Where this comes back

  • Week 33Your own fine-tunes are small-scale SFT, usually with LoRA.

Everything is metered in tokens: roughly 0.75 English words per token, more for code and non-English text. Providers charge separately for input tokens (everything you send: system prompt, history, retrieved documents, tool definitions) and output tokens (what the model generates). Output tokens are usually priced several times higher.

The context window is the maximum input + output tokens in one call. Long windows are convenient, but every token in the window is processed on every call: a 50k-token system prompt is paid for on every turn of a conversation, unless it is cached.

Cost per request = input tokens × input price + output tokens × output price. In an agent loop, the input grows each step because the history is resent, so cost grows roughly quadratically with the number of steps. Budget it explicitly.

cost=nin⋅pin+nout⋅pout,agent cost≈∑s=1S(n0+s⋅Δ) pin\text{cost} = n_{\text{in}}\cdot p_{\text{in}} + n_{\text{out}}\cdot p_{\text{out}},\qquad \text{agent cost} \approx \sum_{s=1}^{S}\big(n_0 + s\cdot\Delta\big)\,p_{\text{in}}

Going deeper

Cost levers at the request level: shorter system prompts, prompt caching for stable prefixes, fewer retrieved chunks, max_tokens caps, smaller models for easy traffic, and batch APIs (typically discounted) for offline work.

Count tokens with the provider's tokenizer or token-counting endpoint, not word counts. Images, tool definitions and system prompts all consume input tokens.

Backend engineer tip: Treat tokens like a metered cloud resource: log them per request, per feature and per user from day one.

Where this comes back

  • Week 28Context engineering is the discipline of spending these tokens well.

At each step the model emits logits over the vocabulary. Temperature divides the logits before softmax: below 1 sharpens the distribution (more deterministic), above 1 flattens it (more random); 0 means greedy. Top-k keeps only the k most likely tokens; top-p (nucleus) keeps the smallest set whose probabilities sum to p.

Guidance: low temperature (0–0.3) for extraction, classification, code and anything you'll parse; moderate (0.7–1.0) for creative writing and brainstorming. Change temperature or top-p, not both at once.

Even at temperature 0, outputs are not perfectly deterministic across calls (batching and floating-point effects). Design for variation: validate outputs, don't assume byte-identical repeats. Reshape a distribution yourself in the simulation.

pi=exp⁡(zi/T)∑jexp⁡(zj/T)p_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}

Going deeper

Structured or constrained decoding masks invalid tokens at every step (for example, anything that would break a JSON Schema), guaranteeing format without retries.

Most chat products default to temperature around 0.7–1.0; APIs used for extraction should set it explicitly rather than inherit a default.

Prefill: all input tokens are processed in parallel in one forward pass, computing and caching each layer's keys and values. It is compute-bound and fast per token. It determines time to first token (TTFT).

Decode: each new token requires a full forward pass for just that position, attending to all cached keys and values. It is sequential and memory-bandwidth-bound. It determines tokens per second, and is why long outputs are slow and why output tokens are priced higher.

The KV cache stores keys and values for every previous token, so each decode step is incremental instead of recomputing the whole context. It grows linearly with context length, often dominating GPU memory at long contexts. Prompt caching reuses the prefill of a shared prefix across requests. Watch both phases in the simulation.

KV bytes≈2⋅nlayers⋅nkv heads⋅dhead⋅ntokens⋅bytes/value\text{KV bytes} \approx 2 \cdot n_{\text{layers}} \cdot n_{\text{kv heads}} \cdot d_{\text{head}} \cdot n_{\text{tokens}} \cdot \text{bytes/value}

Going deeper

Inference arithmetic: decode is memory-bandwidth-bound, since each token reads every weight once. Rough tokens/s per sequence ≈ memory bandwidth ÷ model size in bytes. Batching many sequences amortises those weight reads, which is why throughput and latency trade off.

PagedAttention (vLLM) stores the KV cache in fixed-size blocks like virtual-memory pages, eliminating fragmentation and enabling much larger batches. Continuous batching adds and removes sequences from a running batch at every step.

Where this comes back

  • Week 24The naive generate() loop you wrote recomputed everything; this is the fix.
  • Week 28Prompt caching cuts cost and latency for repeated prefixes.

Rule of thumb: weights take parameters × bytes per parameter. A 7B model needs ~14 GB at 16-bit (bf16), ~7 GB at 8-bit, ~3.5–4 GB at 4-bit. Add the KV cache (grows with context and batch size) and runtime overhead (~10–20%).

Quantisation stores weights in fewer bits (int8, 4-bit formats like GPTQ, AWQ, GGUF's Q4_K_M) with modest quality loss, often small at 8-bit and noticeable but usually acceptable at 4-bit. That is what makes local inference with Ollama or llama.cpp possible on a laptop.

Training needs far more memory than inference: gradients and optimiser states (Adam keeps two extra values per parameter) multiply memory several-fold. That is why QLoRA (4-bit base + small trainable adapters) matters.

VRAMweights≈Nparams×bits8 bytes\text{VRAM}_{\text{weights}} \approx N_{\text{params}} \times \frac{\text{bits}}{8}\ \text{bytes}
text
7B  @ bf16  ≈ 14 GB    @ int8 ≈ 7 GB    @ 4-bit ≈ 3.5–4 GB
70B @ bf16  ≈ 140 GB   @ int8 ≈ 70 GB   @ 4-bit ≈ 35–40 GB

Going deeper

Weight-only quantisation (GPTQ, AWQ, GGUF k-quants) compresses weights and dequantises on the fly; activation quantisation (W8A8, FP8) also speeds up compute. Outlier features in LLM activations are why naive 8-bit quantisation failed and methods like LLM.int8() were needed.

KV-cache quantisation (8-bit or 4-bit keys and values) cuts long-context memory with modest quality impact, which is increasingly the binding constraint.

Where this comes back

  • Week 33QLoRA fine-tunes a 4-bit base model on a single GPU.

Reasoning models are trained to produce long internal chains of thought before answering, trading latency and tokens for accuracy on maths, code and multi-step problems. This is test-time compute: more thinking at inference can substitute for a bigger model. Use them selectively; for simple extraction they're slow and costly.

Hallucination is structural: a language model generates plausible continuations, and nothing in next-token prediction guarantees truth. Models are worse on rare facts, precise numbers, citations and anything after their training cutoff, and they are often confidently wrong.

Engineering mitigations rather than hopes: ground answers in retrieved sources (RAG), require citations, allow and reward 'I don't know', validate structured outputs, and measure with evals.

Going deeper

Reasoning models can be steered with a thinking budget (more tokens on harder problems). Prompting them differs: give the goal and constraints and let them plan, rather than scripting every step.

Hallucination taxonomy: intrinsic (contradicts the provided context) vs extrinsic (unsupported by it), and factuality vs faithfulness. Your RAG evals mostly measure faithfulness.

Where this comes back

  • Week 32Citations and abstaining are the main RAG-side defences.
  • Week 36Evals measure hallucination rates instead of guessing.

In a mixture-of-experts (MoE) transformer, the MLP in each block is replaced by many 'expert' MLPs plus a small router that sends each token to the top-k experts (often 2 of 8, or 8 of hundreds). Only the chosen experts run for that token.

So a model can have a very large total parameter count while each token uses a small active parameter count. Compute per token tracks active parameters; memory tracks total parameters, since all experts must be loaded.

Trade-offs: more memory and more complex serving (expert parallelism, load balancing across experts) in exchange for better quality per unit of compute. Many frontier and open models are MoE; check 'total vs active parameters' on model cards when estimating cost and VRAM.

y=∑i∈TopK(g(x))gi(x) Ei(x)y = \sum_{i \in \text{TopK}(g(x))} g_i(x)\,E_i(x)

Decoding is slow because each token needs a full pass of the big model. Speculative decoding lets a cheap draft (a small model, or extra prediction heads) propose several tokens; the large model checks them all in a single parallel forward pass and accepts the longest correct prefix.

With the right acceptance rule, the output distribution is exactly the large model's. When the draft is usually right (predictable text, code, structured output), speedups of 2–3× are common.

It's mostly a serving-side optimisation (vLLM, TensorRT-LLM and many hosted APIs use it), but it explains why latency can vary with how predictable the output is.

Public signals: benchmark suites (knowledge, reasoning, coding, maths), human-preference arenas, and independent price/speed trackers. They're useful for shortlisting but suffer from contamination (test data leaking into training) and may not resemble your task.

Decide with your eval set: 50–200 representative inputs scored the same way for each candidate, plus cost per 1,000 requests and p95 latency. Also weigh context length, tool-use reliability, structured-output support, data residency, licence (for open weights) and provider reliability.

Revisit regularly: prices drop and new models arrive constantly. Keep a model-agnostic client and an eval harness so switching is a measured, one-afternoon decision.

Where this comes back

  • Week 36Your eval harness turns model choice into a measurement.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Token and cost estimate

    For a chatbot with a 3k-token system prompt, 10 turns averaging 200 input and 300 output tokens, compute cost per conversation with and without prompt caching for two real models' prices.

  2. Core

    VRAM calculator

    Write a function estimating VRAM for weights (by precision) plus KV cache (by layers, KV heads, head dim, context, batch). Check it against a model card or a local Ollama run.

  3. Stretch

    Decoding experiments

    With a local model (Ollama), generate the same prompt at temperatures 0, 0.7 and 1.3 and top-p 0.5/0.95. Describe how diversity and errors change.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 176Monday15 Mar2 h planned

    Pretraining vs SFT vs RLHF / DPO

  2. Day 177Tuesday16 Mar2 h planned

    Tokens, context windows, and cost mechanics

  3. Day 178Wednesday17 Mar2 h planned

    Decoding parameters: temperature, top-p, top-k

  4. Day 179Thursday18 Mar2 h planned

    Inference: prefill vs decode, KV cache, latency drivers

  5. Day 180Friday19 Mar2 h planned

    Quantisation, model sizes, VRAM math

  6. Day 181Saturday20 Mar3 h planned

    Reasoning models, test-time compute, hallucination as a structural property

  7. Day 182Sunday21 MarReview

    Review the week, finish anything unfinished, rest

Watch

Intro to Large Language ModelsPrimary

Andrej Karpathy

Software Is Changing (Again)

Andrej Karpathy · Y Combinator

How might LLMs store facts

3Blue1Brown

Stanford CS25: Transformers United

Stanford Online · playlist

Guest lectures from the people building these models.

Deep Dive into LLMs like ChatGPT

Andrej Karpathy

Large Language Models explained briefly

3Blue1Brown

State of GPT

Andrej Karpathy · Microsoft Build

Pretraining → SFT → RLHF, explained by someone who built it.

RLHF: From Zero to ChatGPT

Hugging Face

LLaMA explained: KV-Cache, RoPE, GQA, SwiGLU

Umar Jamil

Quantization explained with PyTorch

Umar Jamil

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Why are output tokens more expensive and slower than input tokens?
  • Prefill processes input in parallel (compute-bound)
  • Decode generates one token per forward pass (memory-bandwidth-bound)
  • KV cache avoids recomputation but grows with context
How much GPU memory does a 70B model need to serve?
  • Weights: ~140 GB bf16, ~70 GB int8, ~35–40 GB 4-bit
  • Plus KV cache (context × batch) and overhead
  • Multi-GPU or quantisation; MoE: total params set memory
SFT vs RLHF vs DPO?
  • SFT: imitate curated demonstrations
  • RLHF: reward model from preferences + RL with KL penalty
  • DPO: optimise preferences directly, no reward model

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why is generating 1,000 tokens much slower than reading a 1,000-token prompt?

  2. 2.Approximate VRAM for a 7B model's weights at 4-bit?

  3. 3.Best temperature for extracting invoice fields into JSON?

  4. 4.What does the KV cache store?

  5. 5.A base model differs from a chat model mainly because the chat model has had…