The lifecycle: frame the problem (what decision does the prediction drive?), collect and label data, explore, build a baseline, iterate on features and models, evaluate on held-out data, deploy, monitor for drift, and loop back.
Supervised learning maps inputs to known labels: classification (discrete labels) or regression (numbers). Unsupervised learning finds structure without labels: clustering, dimensionality reduction. Self-supervised learning creates labels from the data itself (predict the next word), which is how LLMs are pretrained.
Always start with the dumbest reasonable baseline: predict the majority class, the mean, or yesterday's value. If your model can't beat that clearly, something is wrong.
Going deeper
Google's Rules of ML put it bluntly: rule #1 is 'don't be afraid to launch a product without machine learning'. A heuristic baseline tells you whether the problem is worth a model and gives you the logging infrastructure you'll need anyway.
Most production ML effort goes into data pipelines, monitoring and iteration, not modelling. Plan for data drift (inputs change) and concept drift (the relationship changes) from day one.
Common pitfalls
- Optimising a proxy metric that doesn't move the business decision.
- Skipping the baseline, so you can't tell whether the model adds anything.
Best resources for this lesson
- ArticleRules of Machine Learning (Google) · 43 hard-won lessons from production ML
- CourseIntroduction to ML problem framing (Google)
Where this comes back
- Week 13Your TF-IDF classifier becomes the permanent baseline for every LLM approach.