AI Engineer Path
Phase 1 · Prerequisites12 Oct – 18 Oct

Week 4

Mathematics: linear algebra and calculus

Intuition, not rigour.

Why this week matters

Embeddings are vectors, cosine similarity is a dot product, attention is matrix multiplication, and backpropagation is the chain rule. Four ideas with enormous reach. The Maths track page breaks this into 10 sessions.

Done when

Gradient descent for linear regression written in pure NumPy, and you can explain every line.

Milestone: Gradient descent in pure NumPy

Concepts

7 lessons · tick each one once you could explain it

Geometrically, a vector is an arrow from the origin. Numerically, it is a list of coordinates telling you how much of each basis vector to combine. In 2D the standard basis is i^=(1,0)\hat{i} = (1,0) and j^=(0,1)\hat{j} = (0,1), so (3,2)(3, 2) means three steps along i^\hat{i} plus two along j^\hat{j}.

A linear combination scales and adds vectors: av⃗+bw⃗a\vec{v} + b\vec{w}. The span is every point you can reach that way. If one vector is a multiple of another, they are linearly dependent and add no new direction.

In AI, an embedding is a vector with hundreds of dimensions. You cannot picture 768 dimensions, but the same operations (add, scale, measure angle) work identically.

A basis is a set of ingredients; a vector is a recipe saying how much of each to use.

v⃗=3i^+2j^=[32]\vec{v} = 3\hat{i} + 2\hat{j} = \begin{bmatrix}3\\2\end{bmatrix}

Going deeper

Linear independence and dimension: n independent vectors span an n-dimensional space. Embedding models choose a dimension (384, 768, 1536…) that sets how many independent directions of meaning are available.

Norms measure size: L2 (Euclidean length) is the default; L1 (sum of absolute values) appears in Lasso. Normalising a vector to unit L2 length keeps only its direction.

Best resources for this lesson

Where this comes back

  • Week 12Word2Vec places words as vectors where directions carry meaning.

The single most useful reframe in linear algebra: a matrix is not a grid of numbers, it is a transformation of space that keeps grid lines parallel and evenly spaced and the origin fixed: rotation, scaling, shear, reflection or projection.

Read a 2×2 matrix column by column: the first column is where i^\hat{i} lands, the second is where j^\hat{j} lands. Every other vector follows, because it is a combination of i^\hat{i} and j^\hat{j}. Drag the basis vectors in the simulation and watch the whole grid follow.

Matrix–vector multiplication Ax⃗A\vec{x} is therefore 'apply transformation AA to x⃗\vec{x}'. A neural network layer Wx⃗+b⃗W\vec{x} + \vec{b} is exactly this, followed by a non-linearity.

[abcd][xy]=x[ac]+y[bd]\begin{bmatrix}a & b\\ c & d\end{bmatrix}\begin{bmatrix}x\\y\end{bmatrix} = x\begin{bmatrix}a\\c\end{bmatrix} + y\begin{bmatrix}b\\d\end{bmatrix}

Going deeper

Change of basis: the same vector has different coordinates in different bases, and a matrix of the form P−1APP^{-1}AP represents the same transformation in a new basis. Diagonalisation is choosing the eigenvector basis so the transformation becomes pure scaling.

Non-square matrices map between spaces of different dimension: a 768×3072 weight matrix lifts a 768-d vector into 3072-d, which is exactly what a transformer's MLP does before projecting back down.

Where this comes back

  • Week 15Every dense layer is a matrix transformation plus a non-linearity.

ABAB means 'apply BB, then AA'. That is why the order matters (AB≠BAAB \neq BA in general) and why the row-times-column rule exists: it is what composition works out to. Shapes must chain: (m×n)(n×p)=(m×p)(m\times n)(n\times p) = (m\times p).

The determinant is how much a transformation scales area (2D) or volume (3D). A negative determinant flips orientation; a zero determinant squashes space into a lower dimension, so the transformation can't be undone (no inverse).

Rank is the number of dimensions left after the transformation: the dimension of the column space. A 1000×1000 matrix of rank 8 can be written as the product of a 1000×8 and an 8×1000 matrix. That single fact is the whole idea behind LoRA in week 33.

(AB)ij=∑kAikBkj,det⁡[abcd]=ad−bc(AB)_{ij} = \sum_k A_{ik} B_{kj}, \qquad \det\begin{bmatrix}a&b\\c&d\end{bmatrix} = ad - bc

Going deeper

Cost: multiplying (m×n)(n×p) takes about 2·m·n·p floating-point operations. That number drives every GPU estimate. A forward pass of a transformer costs roughly 2 × parameters FLOPs per token, which you'll use for inference arithmetic in week 26.

Batching turns many matrix–vector products into one matrix–matrix product, which GPUs execute far more efficiently. This is why batch size matters so much for throughput.

Best resources for this lesson

Where this comes back

  • Week 19Attention is a chain of matrix multiplications: QKTQK^T, then times VV.
  • Week 33LoRA learns a low-rank update BABA instead of a full matrix.

Algebraically, the dot product multiplies matching coordinates and sums: a⃗⋅b⃗=∑iaibi\vec{a}\cdot\vec{b} = \sum_i a_i b_i. Geometrically, it equals ∥a⃗∥∥b⃗∥cos⁡θ\|\vec{a}\|\|\vec{b}\|\cos\theta: positive when the vectors point the same way, zero when perpendicular, negative when opposed.

Divide by both lengths and you get cosine similarity, which depends only on the angle. This is the standard measure of semantic similarity between embeddings, because direction encodes meaning and length often encodes things you don't care about (like text length).

If embeddings are already normalised to length 1, cosine similarity is the dot product, which is why many vector databases normalise once at insert time and then use the cheaper dot product.

cos⁡θ=a⃗⋅b⃗∥a⃗∥ ∥b⃗∥=∑iaibi∑iai2 ∑ibi2\cos\theta = \frac{\vec{a}\cdot\vec{b}}{\|\vec{a}\|\,\|\vec{b}\|} = \frac{\sum_i a_i b_i}{\sqrt{\sum_i a_i^2}\,\sqrt{\sum_i b_i^2}}

Going deeper

Because similarity search ranks by dot product, many vector databases store unit-normalised vectors and use 'maximum inner product search'. Some embedding models are trained for raw dot product and encode useful information in vector length, so follow the model card.

The dot product is also a projection: (a⋅b^)b^(a\cdot\hat{b})\hat{b} is the component of aa along bb. Gram–Schmidt orthogonalisation and attention's weighted mixing both build on this.

Where this comes back

  • Week 19Attention scores are dot products between queries and keys.
  • Week 30Vector search ranks documents by cosine or dot-product similarity.

Most vectors get knocked off their span when you apply a matrix. Eigenvectors are the special ones that stay on their own line: Av⃗=λv⃗A\vec{v} = \lambda\vec{v}. The eigenvalue λ\lambda says how much they stretch (or flip, if negative).

They reveal the 'natural axes' of a transformation. PCA (week 10) finds the eigenvectors of the data's covariance matrix: the directions of greatest variance. Repeated application of a matrix is dominated by its largest eigenvalue, which is why gradients in deep RNNs explode or vanish (week 18).

SVD generalises this to any matrix: every matrix is a rotation, then a scaling along axes, then another rotation (UΣVTU\Sigma V^T). Keeping only the largest singular values gives the best low-rank approximation.

Av⃗=λv⃗A=UΣV⊤A\vec{v} = \lambda \vec{v} \qquad A = U\Sigma V^{\top}

Going deeper

Symmetric matrices (like covariance matrices) always have real eigenvalues and orthogonal eigenvectors. That's why PCA's components are perpendicular. The spectral theorem is the formal statement.

Power iteration (multiply a random vector by A repeatedly and normalise) converges to the dominant eigenvector, the same principle behind PageRank. The LoRA simulation uses it to compute a low-rank approximation.

Where this comes back

  • Week 8PCA = eigenvectors of the covariance matrix.
  • Week 18Exploding/vanishing gradients are about eigenvalues above/below 1.

The derivative f′(x)f'(x) is the slope of ff at xx: how much the output changes per tiny change in input. Rules worth knowing cold: power rule ddxxn=nxn−1\frac{d}{dx}x^n = nx^{n-1}, exponentials ddxex=ex\frac{d}{dx}e^x = e^x, and the sigmoid's tidy derivative σ′(x)=σ(x)(1−σ(x))\sigma'(x) = \sigma(x)(1-\sigma(x)).

The chain rule: if y=f(g(x))y = f(g(x)), then dydx=f′(g(x))⋅g′(x)\frac{dy}{dx} = f'(g(x))\cdot g'(x). Wiggle xx; it moves gg by g′g'; that moves ff by f′f' times as much. Rates multiply along the chain.

A neural network is a long chain of functions. Backpropagation is nothing more than applying the chain rule from the loss backwards through every step, reusing intermediate results.

dydx=dydu⋅dudx\frac{dy}{dx} = \frac{dy}{du}\cdot\frac{du}{dx}

Going deeper

For vector-valued functions the derivative becomes the Jacobian matrix, and the chain rule becomes matrix multiplication of Jacobians. Backprop is computing vector–Jacobian products from the output backwards, without ever forming the giant Jacobians.

Numerical differentiation ((f(x+h)−f(x−h))/2h(f(x+h) - f(x-h))/2h) is the standard way to check hand-derived gradients: the 'gradient check' you'll use in week 15.

Where this comes back

  • Week 15Backpropagation is the multivariable chain rule over a computation graph.

For a function of many inputs, a partial derivative measures the slope along one input while holding the others fixed. Stack all of them and you get the gradient ∇L\nabla L, a vector pointing in the direction of steepest increase.

Gradient descent minimises a loss by repeatedly stepping against the gradient: θ←θ−α∇L(θ)\theta \leftarrow \theta - \alpha\nabla L(\theta). The learning rate α\alpha sets the step size: too small and training crawls; too large and it overshoots and diverges. Try both in the simulation.

For linear regression with mean squared error, the gradients have a closed form you can derive by hand, which is this week's capstone.

θt+1=θt−α∇θL(θt),L=1n∑i(wxi+b−yi)2\theta_{t+1} = \theta_t - \alpha \nabla_\theta L(\theta_t), \quad L = \tfrac{1}{n}\sum_i (w x_i + b - y_i)^2
python
import numpy as np
rng = np.random.default_rng(0)
X = rng.uniform(0, 10, 100)
y = 3 * X + 4 + rng.normal(0, 1, 100)

w, b, lr = 0.0, 0.0, 0.01
for step in range(2000):
    pred = w * X + b
    err = pred - y
    loss = (err ** 2).mean()
    dw = 2 * (err * X).mean()   # dL/dw
    db = 2 * err.mean()         # dL/db
    w -= lr * dw
    b -= lr * db
print(w, b)  # ≈ 3, 4

Going deeper

Stochastic gradient descent estimates the gradient from a mini-batch, trading noise for speed; the noise even helps escape sharp minima. The learning-rate stability limit is set by the steepest curvature of the loss (largest Hessian eigenvalue): steps larger than about 2/λ diverge, exactly what the simulation shows.

Feature scaling equalises curvatures across directions, which is why standardised inputs train faster and allow larger learning rates.

Common pitfalls

  • Unscaled features make the loss surface a narrow valley, so gradient descent zig-zags. Standardise inputs (week 7).

Best resources for this lesson

Where this comes back

  • Week 16Momentum, Adam and learning-rate schedules all improve this basic update.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    By hand, then by NumPy

    Compute a 3×3 matrix-vector product, a determinant and a cosine similarity on paper, then verify each with NumPy.

  2. Core

    Capstone A: gradient descent

    Linear regression with gradient descent in pure NumPy: derive the MSE gradients on paper, implement, plot the loss curve, then try three learning rates and explain what you see.

  3. Stretch

    Gradient check

    Implement a numerical gradient check and use it to verify your analytic gradients to 1e-6. Then deliberately introduce a bug and watch the check catch it.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 22Monday12 Oct2 h planned

    3B1B Essence of Linear Algebra ch1-4: vectors, span, linear transformations

  2. Day 23Tuesday13 Oct2 h planned

    3B1B ch5-9: matrix multiplication, determinant, inverse, rank, dot product

  3. Day 24Wednesday14 Oct2 h planned

    3B1B ch10-15: change of basis, eigenvectors, abstract vector spaces

  4. Day 25Thursday15 Oct2 h planned

    Calculus: derivatives, partial derivatives, chain rule

  5. Day 26Friday16 Oct2 h planned

    Gradients and gradient descent intuition

  6. Day 27Saturday17 Oct3 h planned

    MILESTONE: implement gradient descent for linear regression in pure NumPy

  7. Day 28Sunday18 OctReview

    Review the week, finish anything unfinished, rest

Watch

Essence of Linear AlgebraPrimary

3Blue1Brown · playlist

Non-negotiable. About 3 hours; watch every episode.

Essence of CalculusPrimary

3Blue1Brown · playlist

Essential Matrix Algebra for Neural Networks

StatQuest

Cosine Similarity, Clearly Explained

StatQuest

Gradient Descent, Step-by-Step

StatQuest

Maths & Statistics for ML (Data Analysis Process)

CampusX · playlist

Neural Networks

3Blue1Brown · playlist

Gradient descent, how neural networks learn

3Blue1Brown

Backpropagation, intuitively

3Blue1Brown

Backpropagation calculus

3Blue1Brown

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Why is cosine similarity used for embeddings instead of Euclidean distance?
  • Direction carries meaning; magnitude often doesn't
  • On normalised vectors, cosine and L2 give the same ranking
  • Follow what the embedding model was trained for
What does the learning rate do, and what happens if it's too large?
  • Scales each step against the gradient
  • Too small: slow; too large: overshoot and diverge
  • Stability is bounded by the loss surface's steepest curvature
What is the rank of a matrix and why does it matter in ML?
  • Number of independent directions in its column space
  • Low-rank matrices compress into thin factors
  • Basis of PCA and LoRA

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.What do the columns of a 2×2 matrix tell you?

  2. 2.Two normalised embeddings have dot product 0. They are…

  3. 3.A matrix has determinant 0. Which is true?

  4. 4.In gradient descent, why subtract the gradient?

  5. 5.If y = f(g(x)), dy/dx equals…