Concept Lab · Week 4 · Prerequisites
Gradient descent
Roll a ball down a loss surface. Try tiny and huge learning rates.
The idea: Gradients and gradient descent
For a function of many inputs, a partial derivative measures the slope along one input while holding the others fixed. Stack all of them and you get the gradient , a vector pointing in the direction of steepest increase.
Gradient descent minimises a loss by repeatedly stepping against the gradient: . The learning rate sets the step size: too small and training crawls; too large and it overshoots and diverges. Try both in the simulation.
For linear regression with mean squared error, the gradients have a closed form you can derive by hand, which is this week's capstone.
Next simulation: Bayes on a population grid