Concept Lab · Week 10 · Machine Learning
k-means, step by step
Assign, move centroids, repeat, until nothing changes.
The idea: k-means and choosing k
k-means partitions data into k clusters by minimising the within-cluster squared distance. Lloyd's algorithm alternates two steps until nothing changes: assign each point to its nearest centroid, then update each centroid to the mean of its points. Step through it in the simulation.
Choosing k: the elbow method (where adding clusters stops reducing inertia much) and the silhouette score (how much closer points are to their own cluster than the next). Use k-means++ initialisation and several restarts.
Limits: it assumes roughly spherical, similar-sized clusters and needs scaled features. It reappears inside IVF vector indexes (week 30), which cluster embeddings to narrow a search.
Next simulation: Convolution kernels