Concept Lab · Week 9 · Machine Learning
Decision tree splits
Watch a tree carve 2D space into rectangles, one split at a time.
The idea: Decision trees
A tree asks yes/no questions about features (income > 50k?), splitting data into regions with a constant prediction. Each split is chosen greedily to maximise purity gain, measured by Gini impurity or entropy for classification, and variance reduction for regression.
Trees handle mixed feature types, need no scaling, capture interactions and are readable. Alone, they overfit badly: a deep tree memorises the training data. Control it with max_depth, min_samples_leaf or cost-complexity pruning.
Their real power is as building blocks for random forests and gradient boosting (week 10). Watch a tree carve up 2D space in the simulation.
Next simulation: k-means, step by step