AI Engineer Path

Week 11

Advanced concepts + Project 1

Tune, ensemble and explain models, then ship one.

Why this week matters

Explaining a model's behaviour and stating its weaknesses honestly is what makes a project credible. Project 1 is the first portfolio piece.

Done when

Project 1 is deployed behind FastAPI with a README that states its metrics honestly, including where it's weak.

Milestone: Project 1: tabular ML pipeline deployed

Concepts

6 lessons · tick each one once you could explain it

Grid search tries every combination: exhaustive and wasteful. Random search samples combinations and finds good regions faster, because usually only a few hyperparameters matter. Bayesian optimisation (Optuna, TPE) uses past trials to pick promising next ones and prunes bad trials early.

Always tune with cross-validation on training data. Search learning rates and regularisation strengths on a log scale.

Mind the budget: diminishing returns arrive quickly. Better features usually beat more tuning.

python
import optuna
def objective(trial):
    params = {
        "learning_rate": trial.suggest_float("learning_rate", 1e-3, 0.3, log=True),
        "num_leaves": trial.suggest_int("num_leaves", 8, 256, log=True),
        "min_child_samples": trial.suggest_int("min_child_samples", 5, 100),
    }
    return cross_val_score(make_model(**params), X_tr, y_tr, cv=5, scoring="roc_auc").mean()
study = optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Going deeper

Optuna's pruners (median, Hyperband) stop unpromising trials early using intermediate scores, often cutting tuning cost several-fold. Successive halving does the same in scikit-learn (HalvingRandomSearchCV).

Tune the parameters that matter (learning rate, tree size, regularisation) and fix the rest. Log every trial so you can see sensitivity, not just the winner.

Best resources for this lesson

Stacking fits several different base models, collects their out-of-fold predictions, and trains a simple meta-model (often logistic regression) on those predictions. Blending does the same with a single holdout set.

It works when base models make different errors (a linear model, a tree ensemble, a kNN). Gains are usually small and complexity is real, so in production weigh the latency and maintenance cost.

The out-of-fold requirement is leakage prevention again: the meta-model must never see predictions made on data the base model trained on.

Going deeper

Diversity matters more than strength: a stack of three gradient-boosting variants gains little, while boosting + a linear model + kNN can gain more because their errors differ.

Simple weighted averaging of predictions captures most of the benefit with none of the meta-model complexity, so try it first.

Best resources for this lesson

SHAP assigns each feature a contribution to a specific prediction such that contributions sum to the difference between the prediction and the average prediction. It is grounded in Shapley values: a feature's average marginal contribution over all orders of adding features.

Local explanations (a waterfall plot for one prediction) answer 'why this decision?'. Global views (beeswarm, mean |SHAP|) show which features matter overall and in which direction. TreeExplainer is fast and exact for tree ensembles.

Explanations describe the model, not causality in the world, and correlated features share credit in ways that can mislead.

f(x)=E[f(X)]+∑j=1Mϕj(x)f(x) = \mathbb{E}[f(X)] + \sum_{j=1}^{M}\phi_j(x)

Going deeper

KernelSHAP is model-agnostic but slow; TreeSHAP is exact and fast for tree ensembles. SHAP interaction values separate a feature's own effect from its interactions.

SHAP explains the model, not the world: with correlated features, credit can land on either, and 'interventional' vs 'observational' SHAP variants answer different questions.

LIME perturbs an input, watches how the model's prediction changes, and fits a simple interpretable model (like a sparse linear one) locally around that point. It is model-agnostic but can be unstable between runs.

Partial dependence plots (PDP) show the average predicted outcome as one feature varies, marginalising over the others. ICE plots show the same curve per individual row, revealing heterogeneity that averages hide.

Use these to sanity-check models: a PDP showing risk decreasing with more missed payments is a bug, not an insight.

Going deeper

PDPs assume features are independent; with correlated features they average over impossible combinations. Accumulated Local Effects (ALE) plots fix that by averaging local changes instead.

LIME's explanation depends on its sampling neighbourhood and kernel width, which is why explanations can vary between runs.

Save the fitted pipeline (joblib.dump), pinning library versions, since pickles break across versions. Load it once at startup (FastAPI lifespan), validate request bodies with Pydantic, and return predictions with probabilities and a model version.

Ship with: a health endpoint, a Dockerfile, tests for the API contract, and a README that states the metric on held-out data, the baseline it beats, and where the model is weak (segments with worse performance, known failure cases).

Log inputs and predictions so you can detect drift when the live data stops looking like the training data.

Going deeper

Batch vs online: many 'real-time' predictions can be precomputed nightly and looked up, which is cheaper and simpler. Serve online only when inputs arrive at request time.

Shadow deployments (new model scores live traffic without affecting decisions) and canary releases de-risk model updates exactly as they do for code.

Backend engineer tip: This is a normal service. Everything you know about versioning, health checks and contracts applies unchanged.

Where this comes back

  • Week 36LLM serving adds streaming, cost and eval gates to the same foundation.

Without tracking, 'which settings gave that 0.87?' becomes unanswerable within a week. Experiment trackers record each run's parameters, metrics (including curves over time), artefacts (models, plots) and lineage (git commit, data version).

MLflow is open source and self-hostable, with a model registry for promoting versions to staging and production. Weights & Biases is a hosted service with excellent dashboards and sweeps. Both integrate with scikit-learn, PyTorch and Hugging Face in a few lines.

The same habit carries to LLM work: log prompt versions, model versions and eval scores per run so you can compare variants and roll back.

python
import mlflow
mlflow.set_experiment("churn")
with mlflow.start_run():
    mlflow.log_params({"learning_rate": 0.03, "num_leaves": 31})
    model.fit(X_tr, y_tr)
    mlflow.log_metric("val_pr_auc", pr_auc(model, X_val, y_val))
    mlflow.sklearn.log_model(model, "model")

Best resources for this lesson

Where this comes back

  • Week 36LLM observability tools apply the same idea to prompts and traces.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Explain one prediction

    Produce a SHAP waterfall for a single prediction and a beeswarm for the whole model. Write three sentences a non-technical stakeholder would understand.

  2. Core

    Project 1, end to end

    Tune with Optuna (logged to MLflow), explain with SHAP, serve with FastAPI in Docker, and write a README with held-out metrics and known weaknesses.

  3. Stretch

    Shadow deployment

    Add a second model version that scores requests in the background and logs disagreements with the live model.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 71Monday30 Nov2 h planned

    Hyperparameter tuning: grid, random, Optuna

  2. Day 72Tuesday1 Dec2 h planned

    Stacking and blending ensembles

  3. Day 73Wednesday2 Dec2 h planned

    Interpretability: SHAP

  4. Day 74Thursday3 Dec2 h planned

    Interpretability: LIME and partial dependence plots

  5. Day 75Friday4 Dec2 h planned

    PROJECT 1: build the full tabular ML pipeline

  6. Day 76Saturday5 Dec3 h planned

    PROJECT 1: deploy behind FastAPI, write the README

  7. Day 77Sunday6 DecReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Machine LearningPrimary

CampusX · playlist

The roadmap author's own course; node names map to its days.

Project 1

Tabular ML pipeline, deployed

Take a messy real tabular dataset from raw CSV to a prediction served over HTTP, with no manual steps in between.

  • A scikit-learn Pipeline + ColumnTransformer from raw data to prediction
  • A measured improvement from features you engineered yourself
  • Tuned gradient-boosting model with SHAP explanations
  • FastAPI endpoint, Dockerfile, tests
  • README that states metrics honestly, including where the model is weak
Track it on the Projects page

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How would you explain a model's decision to a customer or regulator?
  • Local explanations (SHAP waterfall, LIME) for the specific case
  • Global explanations for overall behaviour
  • Caveats: explanations describe the model, not causation
Walk through deploying a model to production.
  • Package the whole pipeline with pinned versions
  • API with validation, health checks, versioning
  • Monitoring for drift; shadow/canary releases; rollback plan

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why does random search often beat grid search?

  2. 2.Stacking's meta-model must be trained on…

  3. 3.SHAP values for one prediction sum to…

  4. 4.A PDP shows predicted loan risk decreasing as missed payments increase. Most likely…

  5. 5.What belongs in an honest model README?