Mini-batch SGD estimates the gradient on a small batch: noisy but cheap, and the noise even helps escape poor regions. Momentum keeps a running average of past gradients, so steps build speed along consistent directions and damp oscillations across a narrow valley.
RMSprop divides each parameter's step by a running average of its squared gradients: parameters with large, noisy gradients take smaller steps. Adam combines momentum (first moment) with RMSprop-style scaling (second moment), plus bias correction for the early steps.
AdamW decouples weight decay from the adaptive update. It is the default for transformers. Watch all four race across the same awkward valley in the simulation.
Going deeper
Adam's effective step size is roughly the learning rate regardless of gradient scale, which makes it forgiving. Its weakness was weight decay interacting badly with adaptive scaling, fixed by AdamW's decoupled decay.
Newer optimisers (Lion, Sophia, Muon) target memory or speed at LLM scale; AdamW remains the safe default. Optimiser state (two values per parameter for Adam) is a major part of training memory.