AI Engineer Path

Week 10

ML algorithms (2/2)

Ensembles for prediction, and unsupervised methods for structure.

  • Second to cut if behind

Why this week matters

Gradient boosting is still the right answer for most tabular problems in industry. Knowing that keeps you honest when someone proposes an LLM for a spreadsheet.

Unsupervised algorithms are second on the cut list if you fall behind: skim them rather than building.

Done when

You can explain why gradient boosting usually wins on tabular data, and pick a clustering method for a described dataset.

Concepts

6 lessons · tick each one once you could explain it

Bagging (bootstrap aggregating) trains many models on bootstrap samples (random samples with replacement) and averages them. Averaging reduces variance without adding much bias.

A random forest bags decision trees and also considers only a random subset of features at each split, which decorrelates the trees so averaging helps more. The out-of-bag samples (rows a tree didn't see) give a free validation estimate.

Random forests are robust, hard to badly misconfigure and parallelise trivially. They rarely beat well-tuned boosting on accuracy.

Going deeper

Random forests rarely need much tuning: more trees never hurts accuracy (only time), and max_features (features considered per split) is the main knob. Out-of-bag error gives a near-free validation estimate.

ExtraTrees (extremely randomised trees) pick split thresholds at random, faster and sometimes better.

Best resources for this lesson

Boosting builds an ensemble sequentially. AdaBoost re-weights the training points the previous models got wrong. Gradient boosting fits each new tree to the negative gradient of the loss with respect to the current predictions; for squared error, that is simply the residuals.

Each tree's contribution is scaled by a learning rate (shrinkage), so many small corrections add up to a strong model. Trees are kept shallow.

It is gradient descent again, but in 'function space': instead of nudging weights, you add a function that points downhill.

Fm(x)=Fm−1(x)+η hm(x),hm≈−∂L(y,F)∂F∣Fm−1F_m(x) = F_{m-1}(x) + \eta\, h_m(x),\quad h_m \approx -\frac{\partial L(y, F)}{\partial F}\Big|_{F_{m-1}}

Going deeper

Gradient boosting works for any differentiable loss: log loss for classification, quantile loss for intervals, ranking losses (LambdaMART) for search. That generality is why it dominates tabular competitions.

Shrinkage (small learning rate) plus many trees and row/column subsampling is a reliable recipe; early stopping picks the number of trees.

Best resources for this lesson

XGBoost, LightGBM, CatBoost and scikit-learn's HistGradientBoosting* add regularisation, histogram-based splitting (fast on large data), native missing-value handling and categorical support.

The parameters that matter most: learning_rate with n_estimators (use early stopping on a validation set), tree size (max_depth / num_leaves), min_child_samples, and subsampling of rows and columns.

Why they win on tabular data: they handle heterogeneous features, non-linearities and interactions automatically, need little preprocessing and are hard to beat even with deep learning.

python
import lightgbm as lgb
model = lgb.LGBMClassifier(n_estimators=2000, learning_rate=0.03, num_leaves=31,
                           subsample=0.8, colsample_bytree=0.8)
model.fit(X_tr, y_tr, eval_set=[(X_val, y_val)],
          callbacks=[lgb.early_stopping(100)])

Going deeper

LightGBM grows trees leaf-wise (deepest gains first), which is fast and accurate but can overfit small data; limit num_leaves and raise min_child_samples. XGBoost grows depth-wise by default. CatBoost handles categoricals natively with ordered boosting.

Monotonic constraints let you enforce domain knowledge ('risk never decreases as missed payments increase'), which improves trust and robustness.

Where this comes back

  • Week 14Gradient boosting on lag features is usually the production answer for time series too.

k-means partitions data into k clusters by minimising the within-cluster squared distance. Lloyd's algorithm alternates two steps until nothing changes: assign each point to its nearest centroid, then update each centroid to the mean of its points. Step through it in the simulation.

Choosing k: the elbow method (where adding clusters stops reducing inertia much) and the silhouette score (how much closer points are to their own cluster than the next). Use k-means++ initialisation and several restarts.

Limits: it assumes roughly spherical, similar-sized clusters and needs scaled features. It reappears inside IVF vector indexes (week 30), which cluster embeddings to narrow a search.

min⁡C ∑k=1K∑x∈Ck∥x−μk∥2\min_{C}\ \sum_{k=1}^{K}\sum_{x\in C_k}\|x - \mu_k\|^2

Going deeper

k-means minimises within-cluster variance, which is why it prefers spherical, similar-sized clusters. Gaussian mixture models generalise it with elliptical clusters and soft assignments.

Mini-batch k-means scales to millions of points. It's what IVF vector indexes use to build their coarse clusters.

Best resources for this lesson

Where this comes back

  • Week 30IVF indexes use k-means to partition vectors.

Agglomerative hierarchical clustering starts with every point alone and repeatedly merges the closest clusters, producing a dendrogram you can cut at any level. Linkage (single, complete, average, Ward) defines 'closest'.

DBSCAN finds clusters as dense regions: a core point has at least min_samples neighbours within radius eps; clusters grow through connected core points, and isolated points are labelled noise. It finds arbitrarily shaped clusters and doesn't need k, but struggles when densities vary (HDBSCAN fixes that).

Practical use: grouping documents or user queries by embedding to discover topics, as in eval error analysis (week 36).

Going deeper

HDBSCAN extends DBSCAN to clusters of varying density and needs only a minimum cluster size, making it the practical default for clustering embeddings (it's what BERTopic uses).

Clustering has no ground truth, so validate with stability (do clusters persist under resampling?) and by reading samples from each cluster.

Best resources for this lesson

PCA is linear and preserves global variance: good for a first look and for preprocessing. t-SNE preserves local neighbourhoods and reveals clusters well, but distances between clusters and cluster sizes are not meaningful. UMAP is faster, often preserves more global structure, and can transform new points.

These are the standard way to look at embedding spaces: plot your document embeddings coloured by label and you can see whether the embedding model separates what you care about.

Treat the plots as hypotheses, not evidence. Change the perplexity or n_neighbors and the picture can change dramatically.

Going deeper

t-SNE's perplexity sets roughly how many neighbours each point considers; different values can show entirely different pictures of the same data. Run several and trust only structure that persists.

UMAP is fast enough for millions of points and supports transform on new data, so it is the usual choice for embedding dashboards.

Best resources for this lesson

Where this comes back

  • Week 30Visualise embedding spaces before choosing an embedding model.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Forest vs boosting

    Compare a random forest and LightGBM on a tabular dataset; tune LightGBM with early stopping and report the gap.

  2. Core

    Cluster real embeddings

    Embed 2,000 short texts with a sentence-transformer, cluster with k-means and HDBSCAN, and read 5 samples per cluster to judge quality. Visualise with UMAP.

  3. Stretch

    Gradient boosting by hand

    Implement gradient boosting for regression with shallow scikit-learn trees fitted to residuals, and match GradientBoostingRegressor closely.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 64Monday23 Nov2 h planned

    Bagging and Random Forest

  2. Day 65Tuesday24 Nov2 h planned

    Boosting: AdaBoost, Gradient Boosting

  3. Day 66Wednesday25 Nov2 h planned

    XGBoost / LightGBM in practice

  4. Day 67Thursday26 Nov2 h planned

    k-means + choosing k (elbow, silhouette)

  5. Day 68Friday27 Nov2 h planned

    Hierarchical clustering and DBSCAN

  6. Day 69Saturday28 Nov3 h planned

    PCA + t-SNE / UMAP for visualisation

  7. Day 70Sunday29 NovReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Machine LearningPrimary

CampusX · playlist

The roadmap author's own course; node names map to its days.

Machine Learning

StatQuest · playlist

Stanford CS229: Machine Learning (2022)

Stanford Online · playlist

Theory depth. Optional.

PCA, Step-by-Step

StatQuest

Random Forests

StatQuest

AdaBoost, Clearly Explained

StatQuest

Gradient Boost: Regression Main Ideas

StatQuest

XGBoost: Regression

StatQuest

K-means clustering

StatQuest

Hierarchical Clustering

StatQuest

Clustering with DBSCAN

StatQuest

t-SNE, Clearly Explained

StatQuest

UMAP: Main Ideas

StatQuest

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Why does gradient boosting usually win on tabular data?
  • Handles mixed types, non-linearities and interactions
  • Little preprocessing; native missing values
  • Strong regularisation knobs and early stopping
Bagging vs boosting?
  • Bagging: parallel, independent models averaged → reduces variance
  • Boosting: sequential, each fixes predecessors → reduces bias
  • Random forest vs gradient boosting
How would you choose k for k-means?
  • Elbow on inertia, silhouette score
  • Stability across seeds and samples
  • Domain usefulness: read the clusters

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Bagging mainly reduces…

  2. 2.In gradient boosting with squared error, each new tree fits…

  3. 3.Which clustering method labels some points as noise?

  4. 4.In a t-SNE plot, the distance between two clusters…

  5. 5.A colleague proposes an LLM to predict churn from a 40-column customer table. Your first suggestion?