AI Engineer Path
Phase 6 · Build a GPT1 Mar – 7 Mar

Week 24

Let's build GPT

Implement attention from scratch and train a working GPT on your own corpus.

  • Off-roadmap insertion
  • Third to cut if behind

Why this week matters

After this week, 'the model' is no longer a black box: it's code you wrote, and every LLM behaviour you meet later has a mechanical explanation.

Done when

Your GPT generates text, however bad, and you understand every line that produced it.

Milestone: Your own GPT, trained from scratch

Concepts

5 lessons · tick each one once you could explain it

Training data is a long stream of tokens. Sample random chunks of length block_size; the targets are the same chunk shifted by one. A single chunk of 256 tokens therefore holds 256 training examples: predict token 2 from token 1, token 3 from tokens 1–2, and so on.

The causal mask makes this possible in one forward pass: each position may only attend to earlier ones, so no position can cheat by looking at its own target.

Loss is cross-entropy over the vocabulary at every position, averaged. This is the pretraining objective of every GPT-style LLM.

python
def get_batch(data, block_size, batch_size):
    ix = torch.randint(len(data) - block_size, (batch_size,))
    x = torch.stack([data[i:i + block_size] for i in ix])
    y = torch.stack([data[i + 1:i + block_size + 1] for i in ix])  # shifted by one
    return x, y

Going deeper

Teacher forcing during training (always feeding the true previous tokens) creates exposure bias: at generation time the model conditions on its own possibly wrong outputs, so errors compound.

Packing many short documents into fixed-length sequences, separated by end-of-text tokens, keeps GPUs fully utilised during pretraining.

Best resources for this lesson

Karpathy's derivation: the simplest way for a token to use its past is to average previous tokens' vectors. That average can be computed with a lower-triangular matrix multiply. Replace the uniform weights with softmax of masked scores and you have a learned weighted average.

Make the scores data-dependent with queries and keys, aggregate values instead of raw embeddings, scale by dk\sqrt{d_k}, and you have exactly the attention head from week 19, in about 15 lines.

Then: multiple heads, an MLP, residuals and LayerNorm, stacked into blocks.

python
class Head(nn.Module):
    def __init__(self, n_embd, head_size, block_size):
        super().__init__()
        self.key = nn.Linear(n_embd, head_size, bias=False)
        self.query = nn.Linear(n_embd, head_size, bias=False)
        self.value = nn.Linear(n_embd, head_size, bias=False)
        self.register_buffer("tril", torch.tril(torch.ones(block_size, block_size)))
    def forward(self, x):
        B, T, C = x.shape
        k, q, v = self.key(x), self.query(x), self.value(x)
        w = q @ k.transpose(-2, -1) * k.shape[-1] ** -0.5      # (B, T, T)
        w = w.masked_fill(self.tril[:T, :T] == 0, float("-inf"))
        return F.softmax(w, dim=-1) @ v

Going deeper

Batching heads: rather than looping over heads, reshape (B, T, C) to (B, n_head, T, head_size) and run one batched matmul. nanoGPT does exactly this, then uses PyTorch's fused attention kernel.

Best resources for this lesson

The full model: token embedding table plus position embedding table, summed; a stack of blocks; a final LayerNorm; a linear layer mapping each position's vector to logits over the vocabulary.

Scaling knobs: n_embd (width), n_head, n_layer (depth), block_size (context length), dropout. Parameter count is dominated by roughly 12⋅nlayer⋅nembd212 \cdot n_{\text{layer}} \cdot n_{\text{embd}}^2.

Count your model's parameters and compare: GPT-2 small is 124M; your Shakespeare model will be a few million.

params≈12 nlayer d2  +  Vd\text{params} \approx 12\, n_{\text{layer}}\, d^2 \;+\; V d

Going deeper

Weight tying shares the token-embedding matrix with the output projection, saving V·d parameters (38M in GPT-2 small) and often improving quality.

Choose text you can judge: your own writing, a favourite author, code. Encode it with a character or small BPE tokenizer and split 90/10 into train and validation.

Train with AdamW, a learning rate around 3e-4 with warmup, and evaluate both losses every few hundred steps. Validation loss diverging from training loss means overfitting: add dropout, shrink the model or get more data.

Sample periodically. Watching output go from noise to word shapes to plausible phrases is the most convincing demonstration of what next-token prediction learns.

Going deeper

Tiny models on tiny corpora memorise quickly, so watch validation loss and stop early. A useful sanity number: a character-level model on Shakespeare reaches about 1.5 validation loss with a few million parameters.

Best resources for this lesson

Generation is a loop: crop the context to block_size, run the model, take the logits at the last position, apply temperature and optional top-k, softmax, sample one token, append it, repeat.

Only the last position's prediction is used each step, yet the naive loop recomputes attention over the whole context every time. That waste is precisely what the KV cache removes in production inference.

Greedy decoding (always the most likely token) tends to loop and repeat; sampling with a moderate temperature gives more natural text.

python
@torch.no_grad()
def generate(model, idx, max_new, temperature=1.0, top_k=None):
    for _ in range(max_new):
        logits = model(idx[:, -block_size:])[:, -1, :] / temperature
        if top_k:
            v, _ = torch.topk(logits, top_k)
            logits[logits < v[:, [-1]]] = -float("inf")
        nxt = torch.multinomial(F.softmax(logits, dim=-1), 1)
        idx = torch.cat([idx, nxt], dim=1)
    return idx

Going deeper

Repetition and frequency penalties lower the logits of tokens already generated; min-p sampling keeps tokens whose probability is at least a fraction of the top token's, adapting the cutoff to the model's confidence.

Best resources for this lesson

Where this comes back

  • Week 26Temperature, top-p and the KV cache, at production scale.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Train on your own corpus

    Train your GPT on text you know well and save samples at steps 0, 500, 2,000 and the end to watch quality evolve.

  2. Stretch

    Ablate one thing

    Remove positional embeddings, or residual connections, or LayerNorm, retrain and compare validation loss.

  3. Warm-up

    Count parameters

    Compute your model's parameter count by formula and confirm with code.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 162Monday1 Mar2 h planned

    Karpathy 6: WaveNet

  2. Day 163Tuesday2 Mar2 h planned

    Karpathy 7: Let's build GPT (part 1) - attention from scratch

  3. Day 164Wednesday3 Mar2 h planned

    Karpathy 7: Let's build GPT (part 2)

  4. Day 165Thursday4 Mar2 h planned

    Karpathy 7: Let's build GPT (part 3) - finish and train

  5. Day 166Friday5 Mar2 h planned

    MILESTONE: train your GPT on a custom corpus

  6. Day 167Saturday6 Mar3 h planned

    Debug and experiment with hyperparameters

  7. Day 168Sunday7 MarReview

    Review the week, finish anything unfinished, rest

Watch

Neural Networks: Zero to HeroPrimary

Andrej Karpathy · playlist

Type every line. Watching alone is worthless; typing it is transformative.

Let's build GPT: from scratch, in code, spelled outPrimary

Andrej Karpathy

Decoder-Only Transformers

StatQuest

Building makemore Part 5: Building a WaveNet

Andrej Karpathy

Build a Large Language Model (From Scratch)

Sebastian Raschka · playlist

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How does a GPT generate text?
  • Autoregressive: predict the next-token distribution, sample, append, repeat
  • Causal mask so each position sees only the past
  • Decoding settings control randomness
What's the training objective of a GPT, and how is one sequence many examples?
  • Next-token cross-entropy
  • Targets are inputs shifted by one
  • Causal mask lets all positions train in parallel

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.How many training examples does one 256-token chunk provide?

  2. 2.The lower-triangular mask in Karpathy's derivation ensures…

  3. 3.During generation, which logits are used each step?

  4. 4.Validation loss rises while training loss falls. Your GPT is…

  5. 5.Recomputing attention over the whole context every generation step is the inefficiency that…