AI Engineer Path
Phase 6 · Build a GPT22 Feb – 28 Feb

Week 23

micrograd and makemore

Build backpropagation from nothing, then character-level language models.

  • Off-roadmap insertion
  • Third to cut if behind

Why this week matters

Your advantage over a bootcamp graduate is understanding what is beneath an abstraction. Type every line. Watching this is worthless; typing it is transformative.

Off-roadmap insertion. The CampusX roadmap places this on the Research Engineer fork. The case for doing it: when RAG returns garbage, an agent loops or latency spikes at long context, the people who can diagnose it understand tokenisation, attention cost and decoding. The case against: three weeks not spent on the track that pays you. The compromise if you're short: 'Let's build GPT' plus the tokenizer video, roughly four hours.

Done when

Your own micrograd passes a gradient check against PyTorch, and your makemore MLP generates plausible names.

Concepts

5 lessons · tick each one once you could explain it

micrograd wraps a scalar in a Value that stores its data, its gradient, its parent Values and a small _backward function holding its local derivative rule. Overloading +, *, ** and tanh builds the computation graph as you write normal Python.

backward() topologically sorts the graph, sets the output's gradient to 1, and calls each node's _backward in reverse order, accumulating (+=) gradients into parents. That accumulation is the multivariable chain rule's sum over paths.

Then a Neuron, Layer and MLP on top, a training loop, and you have a working neural-network library in about 150 lines. PyTorch is the same idea over tensors.

python
class Value:
    def __init__(self, data, _children=()):
        self.data, self.grad = data, 0.0
        self._prev, self._backward = set(_children), lambda: None
    def __mul__(self, other):
        other = other if isinstance(other, Value) else Value(other)
        out = Value(self.data * other.data, (self, other))
        def _backward():
            self.grad += other.data * out.grad
            other.grad += self.data * out.grad
        out._backward = _backward
        return out
    def backward(self):
        topo, seen = [], set()
        def build(v):
            if v not in seen:
                seen.add(v); [build(c) for c in v._prev]; topo.append(v)
        build(self); self.grad = 1.0
        for v in reversed(topo): v._backward()

Going deeper

Extending micrograd to tensors is the leap to PyTorch: each op stores a backward function over arrays, and broadcasting in the forward pass becomes summation over the broadcast dimensions in the backward pass.

Topological sorting guarantees each node's gradient is complete (all downstream contributions accumulated) before it propagates further. Without it, gradients would be partial.

Common pitfalls

  • Using = instead of += for gradients breaks any node used twice.

Best resources for this lesson

makemore starts with a dataset of names. A bigram model counts how often each character follows each other character, normalises the counts into probabilities, and samples new names one character at a time.

Quality is measured by the negative log-likelihood of the training data: the same cross-entropy loss as before. Then the same model is re-expressed as a single-layer neural network trained by gradient descent, which reaches the same answer as counting.

That equivalence is the key insight: a neural language model is a learned, smoothed, generalising version of counting what comes next.

P(ct∣ct−1)=count(ct−1,ct)∑ccount(ct−1,c),NLL=−1N∑log⁡PP(c_{t}\mid c_{t-1}) = \frac{\text{count}(c_{t-1}, c_t)}{\sum_{c}\text{count}(c_{t-1}, c)},\qquad \text{NLL} = -\frac1N\sum\log P

Going deeper

Smoothing (adding a small count to every pair) is equivalent to L2 regularisation in the neural version: both push the model towards uniform predictions where data is scarce.

Following Bengio et al. (2003): each character gets a learned embedding vector. The model concatenates the embeddings of the last few characters, passes them through a hidden layer and outputs logits over the vocabulary.

Practical lessons: split train/dev/test; find a good learning rate with a sweep; watch for overfitting as the model grows; tune the embedding size and context length.

Embeddings learned this way cluster similar characters (vowels together): the same phenomenon as Word2Vec, discovered by the model on its own.

Going deeper

The Bengio 2003 model already had the core of modern LMs: learned token embeddings, a context window and a softmax over the vocabulary. Transformers changed how context is combined, not the overall shape.

Best resources for this lesson

If the initial loss is far above −log⁡(1/V)-\log(1/V) (the loss of a uniform guess over V characters), the network starts out confidently wrong: shrink the output layer's initial weights.

If tanh activations sit at ±1, they're saturated and their gradients vanish: fix initialisation scale (Kaiming) or normalise. Histograms of activations and gradient-to-data ratios per layer are the diagnostic tools.

BatchNorm is implemented by hand to see why it stabilises training, and why its batch dependence is awkward (and why transformers use LayerNorm instead).

Going deeper

The diagnostics transfer directly to big models: activation histograms per layer, the fraction of saturated units, and gradient-to-data ratios. Large-scale training runs monitor exactly these to catch instabilities early.

Best resources for this lesson

Where this comes back

  • Week 16The theory of initialisation and normalisation, now seen in the numbers.

Derive and code the gradient of every intermediate tensor: cross-entropy, linear layers, BatchNorm, tanh, embeddings. Check each against PyTorch's autograd.

The payoff is fluency with tensor-shaped gradients: why the gradient of a matrix multiply is another matrix multiply with a transpose, and how broadcasting in the forward pass becomes a sum in the backward pass.

It is tedious exactly once, and then you never wonder what loss.backward() is doing again.

Y=XW  ⇒  ∂L∂X=∂L∂YW⊤,∂L∂W=X⊤∂L∂YY = XW \;\Rightarrow\; \frac{\partial L}{\partial X} = \frac{\partial L}{\partial Y}W^\top,\quad \frac{\partial L}{\partial W} = X^\top\frac{\partial L}{\partial Y}

Going deeper

A useful check while deriving: the gradient of any tensor always has exactly that tensor's shape. If your expression produces a different shape, a transpose or a sum over a broadcast dimension is missing.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Type micrograd

    Type micrograd from scratch while watching, then add tanh, exp and division yourself and gradient-check them against PyTorch.

  2. Warm-up

    Bigram by counting and by learning

    Build the bigram model both ways and confirm they reach the same loss.

  3. Stretch

    Manual backprop

    Complete the 'backprop ninja' exercises without looking at the solutions, checking each gradient against autograd.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 155Monday22 Feb2 h planned

    Karpathy 1: micrograd - backprop from scratch (part 1)

  2. Day 156Tuesday23 Feb2 h planned

    Karpathy 1: micrograd (part 2) + exercises

  3. Day 157Wednesday24 Feb2 h planned

    Karpathy 2: makemore - bigram character model

  4. Day 158Thursday25 Feb2 h planned

    Karpathy 3: makemore - MLP

  5. Day 159Friday26 Feb2 h planned

    Karpathy 4: activations, gradients, BatchNorm internals

  6. Day 160Saturday27 Feb3 h planned

    Karpathy 5: becoming a backprop ninja

  7. Day 161Sunday28 FebReview

    Review the week, finish anything unfinished, rest

Watch

Neural Networks: Zero to HeroPrimary

Andrej Karpathy · playlist

Type every line. Watching alone is worthless; typing it is transformative.

Building micrograd

Andrej Karpathy

Building makemore

Andrej Karpathy

Building makemore Part 2: MLP

Andrej Karpathy

Building makemore Part 4: Becoming a Backprop Ninja

Andrej Karpathy

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How does an autograd engine work?
  • Each op records inputs and a local backward rule
  • Topological sort, then reverse traversal from the loss
  • Gradients accumulate across paths

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why must micrograd accumulate gradients with += ?

  2. 2.Initial loss for a 27-character model with a uniform guess is about…

  3. 3.A bigram language model conditions on…

  4. 4.If Y = XW, the gradient of the loss with respect to W is…

  5. 5.Most tanh activations sit at ±1 after initialisation. The problem?