AI Engineer Path

Week 7

Feature engineering (1/2)

Turn raw, messy columns into model-ready features without leaking.

Why this week matters

Feature engineering separates people who have shipped models from people who have watched courses. CampusX's material here is among its strongest.

Done when

A Pipeline + ColumnTransformer that takes raw data to prediction with no manual steps.

Concepts

6 lessons · tick each one once you could explain it

Missingness has causes. MCAR (missing completely at random) is harmless apart from lost data. MAR (missing at random, given other columns) can be modelled from those columns. MNAR (missing not at random: high earners skip the income question) means missingness itself carries signal.

Strategies: drop rows or columns (only when little is lost); impute with mean/median (numeric) or most frequent (categorical); KNN or iterative imputation that predicts missing values from other columns; and always consider adding a missing indicator flag, which lets the model learn from the missingness.

Fit imputers on training data only, inside the pipeline.

python
from sklearn.impute import SimpleImputer
num_imputer = SimpleImputer(strategy="median", add_indicator=True)

Going deeper

Gradient-boosted trees (LightGBM, XGBoost, HistGradientBoosting) handle missing values natively by learning which branch missing values should take, often better than imputation.

Multiple imputation (several plausible imputed datasets) preserves uncertainty for statistical inference; for prediction, a single good imputation plus a missingness flag is usually enough.

Standardisation rescales to mean 0, standard deviation 1. Min-max normalisation squeezes into [0, 1] and is sensitive to outliers. Robust scaling uses the median and IQR, so outliers don't dominate.

Who needs it: kNN, k-means, SVMs, PCA, linear/logistic regression with regularisation, and neural networks, because they use distances or gradients. Tree-based models (random forests, gradient boosting) split on thresholds and are scale-invariant.

In deep learning the same idea reappears inside the network as batch norm and layer norm (week 16).

z=x−μσ,x′=x−xmin⁡xmax⁡−xmin⁡z = \frac{x - \mu}{\sigma}, \qquad x' = \frac{x - x_{\min}}{x_{\max} - x_{\min}}

Going deeper

Scaling changes regularised models' results: L1/L2 penalties treat all coefficients equally, so unscaled features get unequal penalties. Scale before regularised linear models, SVMs and neural nets, always inside the pipeline.

QuantileTransformer maps any distribution to uniform or normal, robust to outliers but distorting distances. Use it when a feature's shape is pathological.

Best resources for this lesson

Where this comes back

  • Week 16Normalisation layers keep activations well-scaled inside networks.

One-hot encoding creates a binary column per category: safe for nominal data but explodes with many categories. Ordinal encoding maps ordered categories (low < medium < high) to integers. Using it on unordered categories invents a false order.

Target encoding replaces each category with the mean target for that category: powerful for high-cardinality features (zip codes, product IDs) but a leakage trap. It must be computed out-of-fold with smoothing.

Handle categories unseen at training time (handle_unknown='ignore') or production will crash on the first new value.

python
from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder, TargetEncoder
OneHotEncoder(handle_unknown="ignore", min_frequency=20)
OrdinalEncoder(categories=[["low", "medium", "high"]])
TargetEncoder(smooth="auto")  # cross-fitted internally

Going deeper

Hashing encoders handle unbounded categories (user agents, URLs) with fixed memory, at the cost of collisions. For deep models, learned embeddings for high-cardinality categories are the standard approach, the same idea as word embeddings.

CatBoost's ordered target statistics are a principled leakage-free target encoding built into the algorithm.

Best resources for this lesson

Where this comes back

  • Week 12Word embeddings are a learned alternative to one-hot encoding words.

Detection: z-scores (for roughly normal data), the IQR rule (beyond 1.5×IQR outside the quartiles), and model-based methods like Isolation Forest for multivariate outliers.

Treatment depends on cause: fix or drop data-entry errors; cap (winsorise) or transform genuine but extreme values; keep them when they are the point (fraud detection is outlier detection).

Robust alternatives sidestep the problem: median instead of mean, MAE instead of RMSE, RobustScaler, tree models.

outlier if x<Q1−1.5 IQR  or  x>Q3+1.5 IQR\text{outlier if } x < Q_1 - 1.5\,\mathrm{IQR} \ \text{ or } \ x > Q_3 + 1.5\,\mathrm{IQR}

Going deeper

Isolation Forest isolates anomalies with random splits: outliers need fewer splits to isolate. Local Outlier Factor compares each point's local density with its neighbours'. Both work for multivariate outliers that per-column rules miss.

In production, outliers in inputs are often a monitoring signal (a broken upstream feed) rather than something to silently clip.

Best resources for this lesson

Many real quantities (income, prices, counts) are right-skewed. A log transform (log1p to handle zeros) compresses the long tail. Box-Cox (positive data) and Yeo-Johnson (any sign) learn the best power transform to make data more normal.

Binning (discretisation) turns a numeric feature into ranges. It can capture non-linear effects for linear models but throws information away.

Transforming the target (e.g. predicting log price) often helps regression. Remember to invert the transform before computing metrics.

python
from sklearn.preprocessing import PowerTransformer, FunctionTransformer
import numpy as np
log = FunctionTransformer(np.log1p, inverse_func=np.expm1)
yj = PowerTransformer(method="yeo-johnson")

Going deeper

Log-transforming the target changes what the model optimises: MSE on log(price) approximates relative error, which is usually what you want for prices and counts. Back-transforming predictions underestimates the mean slightly (smearing correction fixes it).

Splines (SplineTransformer) give linear models smooth non-linear effects without the information loss of hard bins.

Best resources for this lesson

A ColumnTransformer applies different preprocessing to different column groups (numeric vs categorical). A Pipeline chains preprocessing and the model into a single estimator with fit and predict.

Benefits: no leakage (each cross-validation fold refits preprocessing on its own training part), reproducibility, one artefact to deploy, and hyperparameter search over preprocessing choices too (preprocess__num__imputer__strategy).

This is the deliverable for the week and the backbone of Project 1.

python
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.ensemble import HistGradientBoostingClassifier

num = Pipeline([("impute", SimpleImputer(strategy="median")), ("scale", StandardScaler())])
cat = Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
                ("onehot", OneHotEncoder(handle_unknown="ignore"))])
pre = ColumnTransformer([("num", num, num_cols), ("cat", cat, cat_cols)])
model = Pipeline([("pre", pre), ("clf", HistGradientBoostingClassifier())])
model.fit(X_train, y_train)

Going deeper

set_output(transform='pandas') keeps column names through the pipeline, making debugging and SHAP plots far easier. make_column_selector picks columns by dtype so new columns are handled automatically.

Custom transformers subclass BaseEstimator and TransformerMixin with fit and transform, letting domain-specific feature code live inside the same leakage-safe pipeline.

Backend engineer tip: Treat the fitted pipeline like a build artefact: version it, store it, deploy it as one unit.

Best resources for this lesson

Where this comes back

  • Week 11Project 1 deploys this pipeline behind FastAPI.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Imputation shoot-out

    Compare mean, median, KNN and iterative imputation (with and without missing indicators) under 5-fold CV on a dataset with real missing values.

  2. Core

    One pipeline, raw to prediction

    Build a Pipeline + ColumnTransformer that handles numeric, categorical and skewed columns, with no preprocessing outside it. Grid-search one preprocessing choice.

  3. Stretch

    Custom transformer

    Write a scikit-learn-compatible transformer for a domain feature (e.g. ratio features) and slot it into the pipeline with full CV.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 43Monday2 Nov2 h planned

    Missing data: detection and imputation strategies

  2. Day 44Tuesday3 Nov2 h planned

    Scaling: standardisation, normalisation, robust scaling

  3. Day 45Wednesday4 Nov2 h planned

    Categorical encoding: one-hot, ordinal, target encoding

  4. Day 46Thursday5 Nov2 h planned

    Outlier detection and treatment

  5. Day 47Friday6 Nov2 h planned

    Binning, power transforms, skew handling

  6. Day 48Saturday7 Nov3 h planned

    Build a full sklearn Pipeline + ColumnTransformer

  7. Day 49Sunday8 NovReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Machine LearningPrimary

CampusX · playlist

The roadmap author's own course; node names map to its days.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How would you encode a categorical feature with 50,000 unique values?
  • One-hot explodes; consider target encoding (cross-fitted), hashing or learned embeddings
  • Group rare categories
  • Handle unseen categories at inference
Which models need feature scaling and why?
  • Distance- or gradient-based: kNN, SVM, linear with regularisation, neural nets
  • Trees don't: threshold splits are scale-invariant
  • Scale inside the pipeline to avoid leakage

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Which models generally don't need feature scaling?

  2. 2.Ordinal-encoding 'red, green, blue' as 0, 1, 2 for a linear model…

  3. 3.Why put the imputer inside the Pipeline?

  4. 4.Income is heavily right-skewed. A sensible first move for a linear model is…

  5. 5.Target encoding is risky because…