Missingness has causes. MCAR (missing completely at random) is harmless apart from lost data. MAR (missing at random, given other columns) can be modelled from those columns. MNAR (missing not at random: high earners skip the income question) means missingness itself carries signal.
Strategies: drop rows or columns (only when little is lost); impute with mean/median (numeric) or most frequent (categorical); KNN or iterative imputation that predicts missing values from other columns; and always consider adding a missing indicator flag, which lets the model learn from the missingness.
Fit imputers on training data only, inside the pipeline.
from sklearn.impute import SimpleImputer
num_imputer = SimpleImputer(strategy="median", add_indicator=True)Going deeper
Gradient-boosted trees (LightGBM, XGBoost, HistGradientBoosting) handle missing values natively by learning which branch missing values should take, often better than imputation.
Multiple imputation (several plausible imputed datasets) preserves uncertainty for statistical inference; for prediction, a single good imputation plus a missingness flag is usually enough.