In plain words
Linear regression predicts a number as a weighted sum of the inputs, plus a constant: price = a × size + b × rooms + c. It chooses the weights that make the squared prediction errors as small as possible. Each weight says how much the prediction changes per unit of that feature, if everything else stays the same.
Drawing the line through a scatter of points that keeps them as close to it as possible, counting each miss as its distance squared.
In detail
The model predicts . Ordinary least squares picks the weights that minimise mean squared error. There is a closed-form solution (the normal equation), but gradient descent scales better and is what you already built in week 4.
Coefficients are interpretable: holding other features fixed, a one-unit change in changes the prediction by . That interpretation breaks under multicollinearity (correlated features share credit unpredictably).
Assumptions worth checking via residual plots: linearity, constant variance of errors, independent errors.
- the best-fitting weights
- the data matrix: one row per example, one column per feature
- the vector of true target values
- transposed (rows and columns swapped)
- the matrix inverse
Worked example
Fitting a line by hand to four points
- Points (1, 2), (2, 4), (3, 5), (4, 9). Means: x̄ = 2.5, ȳ = 5.
- Slope: .
- Intercept: .
- Predictions: 1.7, 3.9, 6.1, 8.3. Residuals: 0.3, 0.1, −1.1, 0.7, which sum to zero, as least squares guarantees.
- Reading it: each extra unit of x adds 2.2 to the predicted y.
Least squares picks the line that minimises squared residuals; each weight is the effect of one feature, holding the others fixed.
Common mistakes
- Reading coefficients as effects when features are strongly correlated: credit splits between them unpredictably.
- Ignoring residual plots: a curve or a funnel shape in the residuals means the model is missing structure.
Check yourself
A house-price model has weight 3,000 on size_m2. What does that mean?Show answer
Holding the other features fixed, each extra square metre adds 3,000 to the predicted price.
Why do libraries avoid computing directly?Show answer
It's slow for many features and numerically unstable when features are collinear. QR or SVD decompositions, or gradient descent, are more reliable.
Going deeper
The normal equation costs O(d³) and is numerically fragile when features are collinear; libraries use QR or SVD instead. Gradient descent and SGD scale to huge datasets.
Linear models are interpretable only when features are independent-ish and on known scales. Standardise and check variance inflation factors before reading coefficients as 'effects'.
Best resources for this lesson
- InteractiveLinear regression (MLU-Explain)
- BookAn Introduction to Statistical Learning (free book) · chapter 3; the best gentle-but-rigorous ML text
Where this comes back
- Week 4You wrote its gradient descent by hand in the maths capstone.