AI Engineer Path

Concept Lab · Week 4 · Prerequisites

Gradient descent

Roll a ball down a loss surface. Try tiny and huge learning rates.

The idea: Gradients and gradient descent

For a function of many inputs, a partial derivative measures the slope along one input while holding the others fixed. Stack all of them and you get the gradient ∇L\nabla L, a vector pointing in the direction of steepest increase.

Gradient descent minimises a loss by repeatedly stepping against the gradient: θ←θ−α∇L(θ)\theta \leftarrow \theta - \alpha\nabla L(\theta). The learning rate α\alpha sets the step size: too small and training crawls; too large and it overshoots and diverges. Try both in the simulation.

For linear regression with mean squared error, the gradients have a closed form you can derive by hand, which is this week's capstone.

Open the full lesson in week 4

Next simulation: Bayes on a population grid