AI Engineer Path

Concept Lab · Week 10 · Machine Learning

k-means, step by step

Assign, move centroids, repeat, until nothing changes.

The idea: k-means and choosing k

k-means partitions data into k clusters by minimising the within-cluster squared distance. Lloyd's algorithm alternates two steps until nothing changes: assign each point to its nearest centroid, then update each centroid to the mean of its points. Step through it in the simulation.

Choosing k: the elbow method (where adding clusters stops reducing inertia much) and the silhouette score (how much closer points are to their own cluster than the next). Use k-means++ initialisation and several restarts.

Limits: it assumes roughly spherical, similar-sized clusters and needs scaled features. It reappears inside IVF vector indexes (week 30), which cluster embeddings to narrow a search.

Open the full lesson in week 10

Next simulation: Convolution kernels