Concept Lab · Week 15 · Deep Learning
Backpropagation, node by node
Run a forward pass, then watch gradients flow back through a tiny network.
The idea: Backpropagation
The forward pass computes the output and loss, storing intermediate values. The backward pass starts from and walks the graph in reverse. At each node, the local derivative is multiplied by the gradient arriving from above (the chain rule) and passed down to its inputs.
Where a value feeds several later nodes, its gradients from each path are summed. Each node only needs its own local derivative: an add node passes the gradient through unchanged, a multiply node swaps in the other input, ReLU passes it only where the input was positive.
This is why one backward pass costs roughly the same as one forward pass, no matter how many parameters: the gradient for every weight comes out of a single sweep. Step through it node by node in the simulation.
Next simulation: Optimiser race