Maths track
A parallel, start-anytime track that expands weeks 4–5 into daily chunks. Run it alongside whatever week you are on, so the intuition is already built when gradients reappear in deep learning (weeks 15–19).
About 10 focused sessions of 1–1.5 hours, then the two capstones. Compress to 5 if you are moving fast, stretch to 15 as a side channel. Order matters more than speed: don't start statistics before linear algebra and calculus are solid.
Ten sessions
- 1Linear algebra: vectors → matrices as transformations (topics 1–4)
- 2Linear algebra: determinant → dot product (topics 5–7)
- 3Linear algebra: eigenvectors, SVD intuition, review; write 'why cosine similarity = normalised dot product' in your own words
- 4Calculus: derivatives → chain rule (topics 1–3)
- 5Calculus: partial derivatives → gradient (topics 4–5)
- 6Calculus: multivariable chain rule + 3B1B backprop videos (topic 6)
- 7Capstone A: gradient descent for linear regression in pure NumPy
- 8Statistics: descriptive stats → Bayes' theorem (topics 1–3)
- 9Statistics: distributions → hypothesis testing (topics 4–8)
- 10Statistics: A/B testing applied + Capstone B: EDA with 5 non-obvious findings
When the calendar reaches weeks 4–5, treat them as a light review: redo both capstones from memory, faster, and reclaim the time as buffer.
Part 1
Linear Algebra
Embeddings are vectors, cosine similarity is a dot product, attention is matrix multiplication end to end, and LoRA is low-rank matrix decomposition. This is the most-reused maths in the entire path.
Primary: 3Blue1Brown: Essence of Linear Algebra (~3 hrs) · Mathematics for Machine Learning, ch. 2 · CampusX: Maths & Statistics for ML
- 1Vectors3B1B ep 1Not just a list of numbers: magnitude, direction, addition, scalar multiplication.
- 2Linear combinations, span, basis3B1B ep 2Why 'basis' is the right mental model for everything downstream, including embedding spaces.
- 3Matrices as linear transformations3B1B ep 3–4A matrix is a function that moves space (rotate, scale, shear), not a grid of numbers. The highest-leverage reframe in the subject.Try it: Matrices move space
- 4Matrix multiplication3B1B ep 4Composition of transformations. Do row-times-column by hand twice until it is automatic.
- 5Determinant3B1B ep 6How much a transformation scales area or volume, and whether it flips orientation.
- 6Inverse, column space, rank, null space3B1B ep 7, 9When a system has a solution, and why rank-deficient matters (you meet it again with LoRA).
- 7Dot product3B1B ep 9Algebraic (sum of products) and geometric (projection × length). Normalised, it is cosine similarity.Try it: Dot product → cosine similarity
- 8Cross product3B1B ep 10 (skim)Brief and low priority for this path.
- 9Eigenvectors and eigenvalues3B1B ep 14Directions a transformation doesn't rotate, only stretches. Watch twice.Try it: Eigenvectors
- 10SVD (light touch)Supplementary, 20 minAny matrix = rotate, scale, rotate. The intuition behind PCA (week 10) and LoRA (week 33).Try it: LoRA: low-rank updates
Done when
- Explain, without notes, why matrix multiplication is defined the way it is (function composition).
- Compute a 2×2 or 3×3 matrix-vector product and a dot product by hand.
- State what an eigenvector is in one sentence to a non-technical person.
- Explain why cosine similarity is 'the dot product, ignoring magnitude'.
Part 2
Calculus
Backpropagation is the multivariable chain rule applied mechanically across a computation graph. Gradient descent is how every model in this path learns, from week 6 logistic regression to the week 24 GPT.
Primary: 3Blue1Brown: Essence of Calculus (~2 hrs) · Khan Academy: Multivariable Calculus · 3Blue1Brown: Neural Networks ep 3–4 (backprop)
- 1Derivatives as instantaneous rate of change3B1B calc ep 1–2The slope-of-a-curve intuition, not the limit formalism.Try it: The chain rule as a pipeline
- 2Power rule and common derivatives3B1B ep 3Polynomials and exponentials. You need d/dx eˣ constantly for softmax and sigmoid.
- 3Chain rule3B1B ep 4Derivative of a function of a function. Get it mechanically fluent single-variable first.Try it: The chain rule as a pipeline
- 4Partial derivativesKhan AcademyHold every variable constant except one: the natural extension for multi-input functions (every layer).
- 5The gradientKhan AcademyThe vector of all partials; points in the direction of steepest ascent.Try it: Gradient descent
- 6Multivariable chain rule3B1B NN ep 3–4How a change ripples through composed functions. This is backpropagation, full stop.Try it: Backpropagation, node by node
- 7Gradient descent, mechanicallyθ ← θ − α∇L(θ): what the learning rate does and why you subtract.Try it: Gradient descent
- 8Convexity, minima, saddle points (awareness)Enough to know why deep-learning loss landscapes are hard. One paragraph is sufficient.
Done when
- Write gradient descent for linear regression in pure NumPy and explain every line.
Capstone A: gradient descent in pure NumPy
- Initialise weights, compute predictions, compute MSE loss
- Derive by hand, on paper first, the gradient of MSE with respect to weights and bias
- Implement the update loop and watch the loss decrease
- Plot the loss curve
The single most important exercise in the maths track: vectors and calculus fuse into working code, and it is the ancestor of every training loop you will write.
Part 3
Probability & Statistics
Every evaluation decision later (is this RAG change better, is a 3% improvement real, how big must the test set be) is a statistics question in disguise.
Primary: StatQuest: Statistics Fundamentals · CampusX: Maths & Statistics for ML · StatQuest video index · Khan Academy: Statistics & Probability
- 1Descriptive statisticsMean, median, mode, variance, standard deviation, and why variance squares differences (the same reason losses square errors).
- 2Probability basicsSample space, events, independence, conditional probability P(A|B).
- 3Bayes' theoremDerive it from conditional probability; don't memorise it.Try it: Bayes on a population grid
- 4Random variables, PMF vs PDFDiscrete vs continuous, and why P(exactly x) = 0 for a continuous variable.
- 5Key distributionsBernoulli, Binomial, Normal (know cold), Uniform, Poisson (awareness).
- 6Central Limit TheoremWhy sample means tend to normal regardless of the source distribution.Try it: Central Limit Theorem
- 7Sampling and estimationSample vs population, standard error, confidence intervals.
- 8Hypothesis testingNull vs alternative, p-values (and what they don't mean), t-tests.
- 9Correlation vs causationInternalise it well enough to catch it in your own eval analyses.
- 10A/B testing, appliedIs this prompt change really better, or did you get lucky on the eval set? Week 36 revisits it.Try it: Is prompt B really better?
Done when
- Explain why a p-value is not 'the probability the null hypothesis is true'.
- Given 'eval set of 50, new prompt scores 3 points higher', say whether that is enough evidence and what you'd need.
- Explain Bayes' theorem by deriving it, not reciting it.
Part 4
Applied tooling: NumPy, Pandas, plotting
The hands-dirty counterpart to the theory. Do it in parallel with statistics: applying a concept in Pandas the same day you learn it is what makes it stick.
Primary: NumPy for absolute beginners · CampusX: Pandas · 100 Pandas puzzles · Corey Schafer: Matplotlib · CampusX: Seaborn
- 1NumPy broadcastingThe one concept to nail: it makes vectorised gradient descent fast instead of loop-based.
- 2Pandas groupby, merge, missing dataThe three operations behind most EDA.
- 3Plotting a distribution and a correlation matrixEnough Matplotlib/Seaborn to see data without fighting the library.
Done when
- A full EDA on a messy real dataset with five non-obvious findings.
Capstone B: EDA with five non-obvious findings
- Pick a messy real dataset (Kaggle is easiest)
- Clean it and document every decision
- Find five non-obvious findings: a subgroup with a different distribution, a counter-intuitive correlation, a sampling artefact
- Write each finding up with the plot that proves it
'Column X has nulls' is not a finding.