A perceptron computes a weighted sum of inputs plus a bias and passes it through an activation. Alone it can only draw a straight decision boundary, so it can't even learn XOR.
An MLP stacks layers: each layer is a matrix multiply plus bias followed by a non-linearity, . Hidden layers learn intermediate features; the output layer turns them into predictions.
Without non-linearities, any stack of linear layers collapses into a single linear layer (a product of matrices is a matrix). The non-linearity is what buys expressive power. The universal approximation theorem says one wide hidden layer can approximate any continuous function; depth makes it far more efficient.
Going deeper
Depth buys efficiency: some functions that a deep network represents compactly would need exponentially many units in a single hidden layer. Each layer reuses features computed by the one below.
Width and depth both cost parameters, but depth costs latency too, because layers run sequentially. Transformer design is largely a negotiation between the two.
Best resources for this lesson
- InteractiveTensorFlow Playground · train tiny networks in the browser and watch decision boundaries form
- VideoBut what is a neural network? (3Blue1Brown)
- BookDive into Deep Learning: multilayer perceptrons
Where this comes back
- Week 19Every transformer block contains an MLP (the feed-forward layer).