AI Engineer Path

Week 6

ML fundamentals

Frame problems correctly and measure models honestly.

Why this week matters

Metric selection and leakage are the two mistakes that silently invalidate results, in classical ML and in LLM evals alike. Everything you learn here about held-out data comes back in week 36.

Done when

You can name a scenario where accuracy is actively misleading, and say what to use instead.

Concepts

6 lessons · tick each one once you could explain it

The lifecycle: frame the problem (what decision does the prediction drive?), collect and label data, explore, build a baseline, iterate on features and models, evaluate on held-out data, deploy, monitor for drift, and loop back.

Supervised learning maps inputs to known labels: classification (discrete labels) or regression (numbers). Unsupervised learning finds structure without labels: clustering, dimensionality reduction. Self-supervised learning creates labels from the data itself (predict the next word), which is how LLMs are pretrained.

Always start with the dumbest reasonable baseline: predict the majority class, the mean, or yesterday's value. If your model can't beat that clearly, something is wrong.

Going deeper

Google's Rules of ML put it bluntly: rule #1 is 'don't be afraid to launch a product without machine learning'. A heuristic baseline tells you whether the problem is worth a model and gives you the logging infrastructure you'll need anyway.

Most production ML effort goes into data pipelines, monitoring and iteration, not modelling. Plan for data drift (inputs change) and concept drift (the relationship changes) from day one.

Common pitfalls

  • Optimising a proxy metric that doesn't move the business decision.
  • Skipping the baseline, so you can't tell whether the model adds anything.

Best resources for this lesson

Where this comes back

  • Week 13Your TF-IDF classifier becomes the permanent baseline for every LLM approach.

The training set fits model parameters. The validation set chooses hyperparameters and compares models. The test set estimates real-world performance and must be touched once, at the end. Every time you look at test results and change something, the test set leaks into your decisions and your estimate becomes optimistic.

k-fold cross-validation splits data into k folds, trains on k−1 and validates on the remaining fold, k times, then averages. It uses data efficiently and shows how much results vary. Use stratified folds for classification so each fold keeps the class balance, and group folds when several rows belong to the same entity (same user, same patient).

Time-ordered data must be split chronologically: train on the past, validate on the future (week 14).

python
from sklearn.model_selection import StratifiedKFold, cross_val_score
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(pipeline, X_train, y_train, cv=cv, scoring="f1")
print(scores.mean(), scores.std())

Going deeper

Nested cross-validation (an inner loop for tuning, an outer loop for evaluation) gives an unbiased estimate when you also select hyperparameters, important on small datasets where a single split is noisy.

The variance across folds is information: if fold scores range from 0.70 to 0.90, your model choice is less certain than a single average suggests.

Best resources for this lesson

Where this comes back

  • Week 36LLM eval sets need the same discipline: a dev set to iterate on and a held-out set to report.

Leakage produces models that look brilliant in evaluation and fail in production. Common forms: target leakage (a feature derived from the label, like 'refund_issued' when predicting fraud), train–test contamination (fitting a scaler or imputer on the full dataset before splitting), temporal leakage (using future information to predict the past), and duplicate leakage (near-identical rows on both sides of the split).

The fix for preprocessing leakage is structural: put every fitted transformation inside a scikit-learn Pipeline so it is fitted only on training folds.

LLM version: if your eval questions appear in the model's training data, or in the few-shot examples you tuned on, your eval is contaminated.

Going deeper

A practical leakage audit: for every feature, ask 'would I know this value at the moment of prediction, in production?' Timestamps on feature computation make the answer checkable.

Leakage also hides in group structure (the same patient or user in train and test), near-duplicates, and target-derived IDs. GroupKFold and deduplication before splitting fix most of it.

Common pitfalls

  • Scaling or imputing before the split.
  • Random splits on time series.
  • Suspiciously high accuracy: treat it as a bug report.

Best resources for this lesson

Expected test error decomposes into bias² (error from wrong assumptions: a line fit to a curve), variance (error from sensitivity to the particular training sample) and irreducible noise.

As model complexity grows, training error keeps falling but test error falls, bottoms out, then rises: the model starts fitting noise. That U-shape is the central picture of classical ML. Slide the complexity in the simulation and watch it happen.

Remedies for high variance: more data, regularisation, simpler models, bagging, early stopping. For high bias: richer features, more flexible models, boosting. Learning curves (error vs training-set size) tell you which problem you have.

E[(y−f^(x))2]=Bias[f^(x)]2+Var[f^(x)]+σ2\mathbb{E}\big[(y - \hat f(x))^2\big] = \mathrm{Bias}[\hat f(x)]^2 + \mathrm{Var}[\hat f(x)] + \sigma^2

Going deeper

Modern over-parameterised models show double descent: test error falls, rises near the interpolation threshold, then falls again as models get much larger. It doesn't overturn the tradeoff for classical models but explains why huge neural nets generalise.

Learning curves diagnose which side you're on: if train and validation error converge at a high value, you have bias (more data won't help); if a gap persists, you have variance (more data or regularisation will).

Best resources for this lesson

Where this comes back

  • Week 16Dropout, weight decay and early stopping are variance controls for neural nets.

From the confusion matrix: precision = TP/(TP+FP), 'of what I flagged, how much was right?'. Recall = TP/(TP+FN), 'of what was there, how much did I catch?'. F1 is their harmonic mean.

Accuracy is misleading under imbalance: if 1% of transactions are fraud, predicting 'not fraud' always gives 99% accuracy and catches nothing. Use precision/recall, F1 or PR-AUC instead.

ROC-AUC measures ranking quality across all thresholds (true-positive rate vs false-positive rate) and can look rosy under heavy imbalance; PR-AUC is more honest there. The threshold is a business decision: choose it from the cost of a false positive versus a false negative.

Precision=TPTP+FP,Recall=TPTP+FN,F1=2PRP+R\text{Precision} = \frac{TP}{TP+FP},\quad \text{Recall} = \frac{TP}{TP+FN},\quad F_1 = \frac{2PR}{P+R}

Going deeper

Calibration asks whether a predicted 0.8 really means 80% of such cases are positive. Check with a reliability diagram; fix with Platt scaling or isotonic regression. Calibrated scores matter whenever probabilities drive decisions or thresholds.

For multi-class problems, macro-averaging treats each class equally (good for rare classes); micro-averaging weights by frequency. Report per-class metrics whenever classes differ in importance.

Common pitfalls

  • Reporting accuracy on imbalanced data.
  • Tuning the threshold on the test set.

Where this comes back

  • Week 32Retrieval quality uses the same ideas: context precision and context recall in Ragas.

MAE (mean absolute error) is in the target's units and treats all errors linearly: robust to outliers. RMSE squares errors before averaging, so a few large misses dominate: use it when big errors are disproportionately costly. MAPE expresses error as a percentage but explodes near zero.

R² is the fraction of variance explained relative to predicting the mean. It can be high while predictions are still useless for the decision, and it rises with more features even if they are noise (adjusted R² corrects for that).

Pick the metric from the cost of errors, and always look at residual plots: patterns in residuals mean the model is missing structure.

MAE=1n∑∣yi−y^i∣,RMSE=1n∑(yi−y^i)2,R2=1−∑(yi−y^i)2∑(yi−yˉ)2\text{MAE} = \tfrac{1}{n}\sum|y_i - \hat y_i|,\quad \text{RMSE} = \sqrt{\tfrac{1}{n}\sum (y_i - \hat y_i)^2},\quad R^2 = 1 - \frac{\sum (y_i-\hat y_i)^2}{\sum (y_i - \bar y)^2}

Going deeper

MSE is minimised by predicting the conditional mean; MAE by the conditional median. So your loss choice decides what your model estimates. Quantile (pinball) loss predicts any percentile, useful for 'p90 delivery time'.

Always compare against naive baselines (predict the mean; predict yesterday's value). An R² of 0.6 can be worse than 'same as last week' for time series.

Best resources for this lesson

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Make accuracy lie

    Build a dataset with 2% positives, train a dumb majority classifier and a real model, and compare accuracy, precision, recall, F1, ROC-AUC and PR-AUC for both.

  2. Core

    Catch a leak

    Train a model with scaling fitted before the split and another with a Pipeline. Then add a target-derived feature. Measure how much each leak inflates the score.

  3. Stretch

    Learning curves

    Plot learning curves for a shallow and a deep tree on the same data and diagnose bias vs variance from the plots alone.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 36Monday26 Oct2 h planned

    ML lifecycle, problem framing, supervised vs unsupervised

  2. Day 37Tuesday27 Oct2 h planned

    Train / validation / test splits, cross-validation, data leakage

  3. Day 38Wednesday28 Oct2 h planned

    Bias-variance tradeoff, overfitting and underfitting

  4. Day 39Thursday29 Oct2 h planned

    Classification metrics: accuracy, precision, recall, F1, ROC-AUC, PR-AUC

  5. Day 40Friday30 Oct2 h planned

    Regression metrics: RMSE, MAE, R2 - and when each one lies

  6. Day 41Saturday31 Oct3 h planned

    CampusX catch-up + build your first sklearn pipeline

  7. Day 42Sunday1 NovReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Machine LearningPrimary

CampusX · playlist

The roadmap author's own course; node names map to its days.

Machine Learning

StatQuest · playlist

Machine Learning Specialization

Andrew Ng · DeepLearning.AI · playlist

Bias and Variance

StatQuest

Cross Validation

StatQuest

ROC and AUC, Clearly Explained

StatQuest

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Describe three ways data leakage happens and how to prevent each.
  • Preprocessing fitted on all data → pipelines
  • Target-derived or future features → 'available at prediction time?' audit
  • Group/temporal overlap → GroupKFold, chronological splits
When is ROC-AUC misleading, and what would you use instead?
  • Heavy class imbalance makes FPR look small
  • PR-AUC or precision at k focus on the positive class
  • Pick the threshold from business costs
What's the bias–variance tradeoff?
  • Error = bias² + variance + noise
  • Simple models underfit; flexible models overfit
  • Diagnose with train/validation gaps and learning curves

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.1% of emails are spam. A model labels everything 'not spam'. Its accuracy is…

  2. 2.You fit a StandardScaler on the full dataset, then split. What went wrong?

  3. 3.Training error keeps falling while validation error rises. This is…

  4. 4.Missing a fraud case is far costlier than a false alarm. Prioritise…

  5. 5.When should you touch the test set?