Grid search tries every combination: exhaustive and wasteful. Random search samples combinations and finds good regions faster, because usually only a few hyperparameters matter. Bayesian optimisation (Optuna, TPE) uses past trials to pick promising next ones and prunes bad trials early.
Always tune with cross-validation on training data. Search learning rates and regularisation strengths on a log scale.
Mind the budget: diminishing returns arrive quickly. Better features usually beat more tuning.
import optuna
def objective(trial):
params = {
"learning_rate": trial.suggest_float("learning_rate", 1e-3, 0.3, log=True),
"num_leaves": trial.suggest_int("num_leaves", 8, 256, log=True),
"min_child_samples": trial.suggest_int("min_child_samples", 5, 100),
}
return cross_val_score(make_model(**params), X_tr, y_tr, cv=5, scoring="roc_auc").mean()
study = optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)Going deeper
Optuna's pruners (median, Hyperband) stop unpromising trials early using intermediate scores, often cutting tuning cost several-fold. Successive halving does the same in scikit-learn (HalvingRandomSearchCV).
Tune the parameters that matter (learning rate, tree size, regularisation) and fix the rest. Log every trial so you can see sensitivity, not just the winner.