AI Engineer Path

Week 8

Feature engineering (2/2)

Create, select and compress features, and handle imbalance.

Why this week matters

Good features beat clever models on tabular data. You will also meet PCA, the first appearance of the 'low-rank' idea that returns as LoRA.

Done when

A measured improvement from features you engineered yourself, on last week's pipeline.

Concepts

5 lessons · tick each one once you could explain it

Ratios (debt / income), differences (days since last purchase), aggregates per entity (a customer's average order value), interactions (price × quantity) and flags (is_weekend) often add more than any algorithm change.

Domain knowledge tells you which ones matter. Ask: what would a human expert look at to make this decision?

Measure every feature honestly: keep a fixed CV setup and compare scores with and without it. Aggregates per entity must be computed without peeking at the target row or the future.

Going deeper

Time-aware aggregates (a customer's spend in the previous 30 days, computed as of each prediction date) are the most powerful features in transactional data, and the most leakage-prone. Build them with explicit 'as of' timestamps.

Automated feature tools (Featuretools) generate candidates by stacking aggregations over relationships, but domain-driven features still tend to win.

Common pitfalls

  • Aggregates that include the current row's label (leakage).
  • Adding dozens of features without measuring each one's contribution.

Filter methods score features independently of a model: correlation, mutual information, chi-squared. Fast, but blind to interactions. Wrapper methods search subsets using a model (recursive feature elimination). Accurate but expensive. Embedded methods select during training: L1 regularisation zeroes out weights; tree importances rank splits.

Use permutation importance on validation data rather than default tree importances, which favour high-cardinality features.

Do selection inside cross-validation, or the selection step itself leaks.

python
from sklearn.inspection import permutation_importance
r = permutation_importance(model, X_val, y_val, n_repeats=10, random_state=0)
ranked = sorted(zip(r.importances_mean, X_val.columns), reverse=True)

Going deeper

Correlated features split importance between them, so each looks less important than the pair is. Cluster correlated features and evaluate groups, or drop one per cluster.

Boruta and SHAP-based selection compare real features against shuffled 'shadow' copies: a principled way to keep only features that beat noise.

PCA finds orthogonal directions (principal components) ordered by how much variance they capture. They are the eigenvectors of the covariance matrix, or equivalently come from the SVD of the centred data matrix.

Projecting onto the top k components compresses features while keeping most of the information. The explained variance ratio tells you how much each component keeps. Standardise first, or large-scale features dominate.

Uses: visualisation, de-noising, speeding up models, removing multicollinearity. Cost: components are combinations of features and are hard to interpret.

Xc=UΣV⊤,Z=XcVk  (top-k components)X_c = U\Sigma V^\top,\qquad Z = X_c V_k \ \ (\text{top-}k\text{ components})

Going deeper

PCA via SVD: centre the data matrix X, compute X=UΣV⊤X = U\Sigma V^\top; the columns of V are the principal directions and σi2/(n−1)\sigma_i^2/(n-1) their variances. Truncated or randomised SVD makes it fast on large, sparse data (it's called LSA when applied to TF-IDF).

Kernel PCA captures non-linear structure; for visualisation, t-SNE and UMAP usually show clusters better, but PCA remains the honest, linear first look.

Where this comes back

  • Week 4Eigenvectors and SVD are the machinery underneath.
  • Week 33LoRA uses the same low-rank idea to fine-tune cheaply.

First, use metrics that reflect imbalance (precision/recall, PR-AUC). Then options: class weights (class_weight='balanced') make errors on the minority class cost more; undersampling the majority; oversampling the minority; SMOTE synthesises minority examples by interpolating between neighbours.

Often the simplest win is threshold tuning: keep the model, move the decision threshold to the point that matches your precision/recall needs.

Resample only the training data, inside CV. Resampling before splitting leaks synthetic copies of test points into training.

Going deeper

Class weights and resampling distort predicted probabilities: the model's scores no longer match real-world frequencies. If you need calibrated probabilities, recalibrate on untouched validation data afterwards.

For extreme imbalance (1 in 10,000), frame it as anomaly detection or ranking, and evaluate with precision at k (how many of the top-k flagged cases are real).

Common pitfalls

  • SMOTE before the train/test split.
  • Judging an imbalanced model by accuracy.

Best resources for this lesson

From timestamps: hour, day of week, month, is_holiday, time since an event. Encode cyclical features with sine/cosine so 23:00 and 00:00 are close.

From short text fields: length, word count, presence of keywords, TF-IDF vectors (week 12) fed into a tabular model alongside other features.

These are the bridges from tabular ML to the NLP and time-series weeks.

hoursin⁡=sin⁡ ⁣(2π⋅h24),hourcos⁡=cos⁡ ⁣(2π⋅h24)\text{hour}_{\sin} = \sin\!\left(\frac{2\pi \cdot h}{24}\right),\quad \text{hour}_{\cos} = \cos\!\left(\frac{2\pi \cdot h}{24}\right)

Going deeper

Holiday calendars, paydays and promotions are often the strongest temporal signals. Encode them explicitly rather than hoping the model infers them.

For short text in tabular data, a small sentence-embedding model can replace TF-IDF columns, and the embedding dimensions feed straight into gradient boosting.

Where this comes back

  • Week 12TF-IDF turns text into numeric features properly.
  • Week 14Lag features and chronological splits for time series.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    PCA on images

    Run PCA on the digits dataset, plot explained variance, and reconstruct images from 5, 20 and 50 components.

  2. Core

    Feature engineering that pays

    Add 5 domain features to last week's pipeline and measure each one's contribution with the same CV setup. Keep only those that help.

  3. Stretch

    Imbalance strategies compared

    Compare class weights, SMOTE (inside the pipeline) and threshold tuning on an imbalanced dataset using PR-AUC.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 50Monday9 Nov2 h planned

    Feature construction and domain-driven features

  2. Day 51Tuesday10 Nov2 h planned

    Feature selection: filter, wrapper, embedded methods

  3. Day 52Wednesday11 Nov2 h planned

    Dimensionality reduction with PCA

  4. Day 53Thursday12 Nov2 h planned

    Imbalanced data: resampling, SMOTE, class weights

  5. Day 54Friday13 Nov2 h planned

    Datetime and text-derived features for tabular models

  6. Day 55Saturday14 Nov3 h planned

    Refactor last week's pipeline with new features; measure the delta

  7. Day 56Sunday15 NovReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Machine LearningPrimary

CampusX · playlist

The roadmap author's own course; node names map to its days.

PCA, Step-by-Step

StatQuest

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How do you handle a dataset with 1% positive class?
  • Right metrics first: precision/recall, PR-AUC
  • Class weights, resampling inside CV, threshold tuning
  • Consider anomaly-detection framing for extreme cases
What does PCA do and when would you use it?
  • Projects onto orthogonal directions of max variance (eigenvectors/SVD)
  • Compression, visualisation, de-noising, collinearity
  • Loses interpretability; standardise first

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.PCA components are…

  2. 2.Where should SMOTE be applied?

  3. 3.Why encode hour-of-day with sine and cosine?

  4. 4.Default tree feature importances can mislead because they…

  5. 5.Cheapest first fix when a classifier's recall is too low on a rare class?