AI Engineer Path

Concept Lab · Week 16 · Deep Learning

Optimiser race

SGD, Momentum, RMSprop and Adam race across the same awkward valley.

The idea: SGD, momentum, RMSprop and Adam

Mini-batch SGD estimates the gradient on a small batch: noisy but cheap, and the noise even helps escape poor regions. Momentum keeps a running average of past gradients, so steps build speed along consistent directions and damp oscillations across a narrow valley.

RMSprop divides each parameter's step by a running average of its squared gradients: parameters with large, noisy gradients take smaller steps. Adam combines momentum (first moment) with RMSprop-style scaling (second moment), plus bias correction for the early steps.

AdamW decouples weight decay from the adaptive update. It is the default for transformers. Watch all four race across the same awkward valley in the simulation.

Open the full lesson in week 16

Next simulation: Self-attention, Q·K·V