AI Engineer Path
Phase 4 · Deep Learning28 Dec – 3 Jan

Week 15

Foundations of neural networks

Understand a neural network all the way down: forward pass, loss, backward pass, update.

Why this week matters

Every model from here to the end of the path, GPT included, is this loop at larger scale.

Done when

A 2-layer network written in NumPy, trained and debugged by you.

Concepts

5 lessons · tick each one once you could explain it

A perceptron computes a weighted sum of inputs plus a bias and passes it through an activation. Alone it can only draw a straight decision boundary, so it can't even learn XOR.

An MLP stacks layers: each layer is a matrix multiply plus bias followed by a non-linearity, h=ϕ(Wx+b)h = \phi(Wx + b). Hidden layers learn intermediate features; the output layer turns them into predictions.

Without non-linearities, any stack of linear layers collapses into a single linear layer (a product of matrices is a matrix). The non-linearity is what buys expressive power. The universal approximation theorem says one wide hidden layer can approximate any continuous function; depth makes it far more efficient.

h(1)=ϕ(W(1)x+b(1)),y^=W(2)h(1)+b(2)h^{(1)} = \phi(W^{(1)}x + b^{(1)}),\qquad \hat y = W^{(2)}h^{(1)} + b^{(2)}

Going deeper

Depth buys efficiency: some functions that a deep network represents compactly would need exponentially many units in a single hidden layer. Each layer reuses features computed by the one below.

Width and depth both cost parameters, but depth costs latency too, because layers run sequentially. Transformer design is largely a negotiation between the two.

Best resources for this lesson

Where this comes back

  • Week 19Every transformer block contains an MLP (the feed-forward layer).

Sigmoid squashes to (0, 1) and tanh to (−1, 1); both saturate (flat gradients at the extremes), which slows learning in deep networks. ReLU (max⁡(0,x)\max(0, x)) is cheap and doesn't saturate for positive inputs, which made deep networks trainable. Units can 'die' if they get stuck negative.

GELU is a smooth ReLU-like curve used in BERT and GPT. SwiGLU variants appear in modern LLMs like LLaMA.

Softmax turns a vector of scores (logits) into a probability distribution: exponentiate, then normalise. It is the output of every classifier and every LLM's next-token prediction.

ReLU(x)=max⁡(0,x),softmax(z)i=ezi∑jezj\text{ReLU}(x) = \max(0,x),\qquad \text{softmax}(z)_i = \frac{e^{z_i}}{\sum_j e^{z_j}}

Going deeper

Gated activations (GLU, SwiGLU) multiply one linear projection by an activated second projection, giving the network a learned 'gate'. Most modern LLMs use SwiGLU in their MLP blocks.

Softmax is invariant to adding a constant to all logits, which is why implementations subtract the max before exponentiating for numerical stability.

Where this comes back

  • Week 26Temperature divides the logits before softmax, reshaping the distribution.

Regression uses MSE or MAE. Classification uses cross-entropy: the negative log of the probability the model assigned to the correct class. Assigning 0.9 costs 0.105; assigning 0.01 costs 4.6. Confident wrong answers are punished hard.

Softmax followed by cross-entropy has a beautifully simple gradient with respect to the logits: predicted probabilities minus the one-hot target. Frameworks fuse them (nn.CrossEntropyLoss takes raw logits) for numerical stability.

An LLM's pretraining loss is exactly this, averaged over every next-token prediction in the training data. Perplexity is just elosse^{\text{loss}}.

L=−∑kyklog⁡p^k=−log⁡p^correct,∂L∂z=p^−yL = -\sum_k y_k \log \hat p_k = -\log \hat p_{\text{correct}},\qquad \frac{\partial L}{\partial z} = \hat p - y

Going deeper

Cross-entropy is maximum-likelihood estimation in disguise: minimising it maximises the probability the model assigns to the observed data. Label smoothing (target 0.9 instead of 1.0) curbs overconfidence.

Information-theory view: cross-entropy in bits is the average number of bits needed to encode the true outcome using the model's distribution, which is why LLM losses are sometimes reported as bits per byte.

Best resources for this lesson

The forward pass computes the output and loss, storing intermediate values. The backward pass starts from ∂L/∂L=1\partial L/\partial L = 1 and walks the graph in reverse. At each node, the local derivative is multiplied by the gradient arriving from above (the chain rule) and passed down to its inputs.

Where a value feeds several later nodes, its gradients from each path are summed. Each node only needs its own local derivative: an add node passes the gradient through unchanged, a multiply node swaps in the other input, ReLU passes it only where the input was positive.

This is why one backward pass costs roughly the same as one forward pass, no matter how many parameters: the gradient for every weight comes out of a single sweep. Step through it node by node in the simulation.

∂L∂w=∂L∂y^⋅∂y^∂h⋅∂h∂w\frac{\partial L}{\partial w} = \frac{\partial L}{\partial \hat y}\cdot\frac{\partial \hat y}{\partial h}\cdot\frac{\partial h}{\partial w}

Going deeper

Backprop is reverse-mode automatic differentiation: one backward pass gives the gradient with respect to all parameters at roughly the cost of a forward pass. Forward-mode would need one pass per parameter.

Memory, not compute, is often the constraint: the backward pass needs stored activations from the forward pass. Gradient checkpointing recomputes some activations to save memory, trading compute for capacity.

Where this comes back

  • Week 4The chain rule you learned there, now applied to a whole graph.
  • Week 23You'll build an autograd engine that does this automatically (micrograd).

Build it on a toy dataset (spirals or two moons): input → hidden layer with ReLU → output with softmax. Initialise small random weights, compute the forward pass, cross-entropy loss, gradients by hand, update with plain gradient descent.

Gradient checking catches mistakes: compare your analytic gradient for a few weights against a numerical estimate (L(w+ϵ)−L(w−ϵ))/2ϵ(L(w+\epsilon) - L(w-\epsilon))/2\epsilon. They should agree to several decimal places.

Debugging habits worth keeping: overfit a tiny batch first (the loss should go to near zero); plot the loss curve; check shapes at every step.

python
# X: (N, 2), y: (N,) integer labels, K classes
W1 = 0.01 * rng.standard_normal((2, 64)); b1 = np.zeros(64)
W2 = 0.01 * rng.standard_normal((64, K)); b2 = np.zeros(K)
for step in range(5000):
    h = np.maximum(0, X @ W1 + b1)                  # forward
    logits = h @ W2 + b2
    p = np.exp(logits - logits.max(1, keepdims=True)); p /= p.sum(1, keepdims=True)
    loss = -np.log(p[np.arange(N), y]).mean()
    d = p.copy(); d[np.arange(N), y] -= 1; d /= N  # dL/dlogits
    dW2 = h.T @ d; db2 = d.sum(0)
    dh = d @ W2.T; dh[h <= 0] = 0                    # back through ReLU
    dW1 = X.T @ dh; db1 = dh.sum(0)
    for P, G in ((W1, dW1), (b1, db1), (W2, dW2), (b2, db2)):
        P -= 1.0 * G

Going deeper

Karpathy's training recipe applies even here: become one with the data, get an end-to-end skeleton with a dumb baseline working, overfit a single batch, then regularise, then tune.

Common silent bugs: forgetting to shuffle, labels misaligned with inputs, softmax applied twice, learning rate off by 10×. Visualise predictions, not just the loss.

Common pitfalls

  • Forgetting to subtract the max before exponentiating (overflow).
  • Initialising all weights to zero: every hidden unit learns the same thing.

Best resources for this lesson

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Playground intuition

    In TensorFlow Playground, find the smallest network that separates the spiral dataset. Note what changes when you remove the non-linearity.

  2. Core

    2-layer net in NumPy

    Implement forward, softmax cross-entropy, backward and SGD for a 2-layer net on spirals. Gradient-check it. Reach >95% training accuracy.

  3. Stretch

    Add a layer, generally

    Refactor into Layer classes with forward/backward so you can stack any number of layers, the shape of every framework.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 99Monday28 Dec2 h planned

    Perceptron and MLP architecture

  2. Day 100Tuesday29 Dec2 h planned

    Forward pass; activations: ReLU, GELU, sigmoid, softmax

  3. Day 101Wednesday30 Dec2 h planned

    Loss functions and cross-entropy

  4. Day 102Thursday31 Dec2 h planned

    Backpropagation worked through by hand

  5. Day 103Friday1 Jan2 h planned

    Implement a 2-layer neural network in pure NumPy

  6. Day 104Saturday2 Jan3 h planned

    Train it on a toy dataset and debug it

  7. Day 105Sunday3 JanReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Deep LearningPrimary

CampusX · playlist

Neural Networks

3Blue1Brown · playlist

But what is a neural network?

3Blue1Brown

Gradient descent, how neural networks learn

3Blue1Brown

Backpropagation, intuitively

3Blue1Brown

Backpropagation calculus

3Blue1Brown

Neural Networks / Deep Learning

StatQuest · playlist

Backpropagation Main Ideas

StatQuest

ReLU In Action

StatQuest

ArgMax and SoftMax

StatQuest

Cross Entropy

StatQuest

Neural Networks and Deep Learning (DLS Course 1)

Andrew Ng · DeepLearning.AI · playlist

NYU Deep Learning (LeCun & Canziani)

Alfredo Canziani · playlist

CMU 11-785 Introduction to Deep Learning

Carnegie Mellon University · playlist

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Walk through backpropagation for a 2-layer network.
  • Forward pass stores activations
  • Gradient of loss w.r.t. logits = p − y for softmax+CE
  • Chain rule backwards through each layer: dW = activationsᵀ · upstream
  • Update all weights with the gradients
Why do we need non-linear activations?
  • Without them, stacked layers collapse to one linear map
  • Non-linearity lets networks approximate complex functions
  • ReLU/GELU avoid saturation of sigmoid/tanh

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why do we need non-linear activations between layers?

  2. 2.The gradient of softmax + cross-entropy with respect to the logits is…

  3. 3.A value feeds two later nodes. During backprop its gradient is…

  4. 4.Best first debugging step for a new network?

  5. 5.Perplexity relates to cross-entropy loss how?