AI Engineer Path
Phase 1 · Prerequisites19 Oct – 25 Oct

Week 5

Statistics and the data stack

Probabilistic literacy plus working NumPy and Pandas.

Why this week matters

Every evaluation decision later (sample size for an eval set, whether a 3% improvement is real) is a statistics question.

Done when

A full EDA on a messy real dataset, with five non-obvious findings.

Milestone: EDA with five non-obvious findings

Concepts

8 lessons · tick each one once you could explain it

A random variable maps outcomes to numbers. Discrete ones have a probability mass function (PMF); continuous ones a probability density function (PDF), where only intervals have non-zero probability: P(X=1.5)=0P(X = 1.5) = 0, but P(1<X<2)P(1 < X < 2) is the area under the curve.

The expectation E[X]E[X] is the probability-weighted average. Variance is the expected squared distance from the mean; its square root is the standard deviation. Squaring penalises large deviations heavily, for exactly the same reason mean squared error does.

Distributions to know: Bernoulli (one yes/no trial: is this answer correct?), Binomial (count of successes in n trials: correct answers on an eval set), Normal (sums and averages of many things, know it cold), Uniform, and Poisson for counts per interval.

E[X]=∑xx P(x),Var(X)=E[(X−E[X])2]E[X] = \sum_x x\,P(x), \qquad \mathrm{Var}(X) = E\big[(X - E[X])^2\big]

Going deeper

The variance of a sum of independent variables is the sum of their variances, which is why averaging n samples shrinks variance by 1/n. Correlated samples (for example, several eval questions generated from the same document) shrink it much less.

Heavy-tailed distributions (latencies, costs, token counts) make the mean misleading. Report medians and high percentiles (p95, p99) for anything operational.

Best resources for this lesson

Where this comes back

  • Week 36Eval accuracy is a binomial proportion with a standard error you can compute.

P(A∣B)P(A\mid B) is the probability of A given that B happened: restrict the world to B, then ask how much of it is A. By definition P(A∣B)=P(A∩B)/P(B)P(A\mid B) = P(A \cap B)/P(B).

Write the joint probability both ways, P(A∩B)=P(A∣B)P(B)=P(B∣A)P(A)P(A\cap B) = P(A\mid B)P(B) = P(B\mid A)P(A), and divide: that is Bayes' theorem. It flips a conditional around: from 'how likely is this evidence if the hypothesis is true' to 'how likely is the hypothesis given the evidence'.

The classic trap is ignoring the base rate. A 99%-accurate detector for something that affects 1 in 1,000 people still produces mostly false positives. The simulation lets you see it as a population grid.

P(H∣E)=P(E∣H) P(H)P(E)P(H\mid E) = \frac{P(E\mid H)\,P(H)}{P(E)}

Going deeper

In odds form, Bayes is easy to do in your head: posterior odds = prior odds × likelihood ratio. A test with a 99:1 likelihood ratio applied to 1:999 prior odds gives 99:999, about 9%.

Bayesian updating is sequential: today's posterior is tomorrow's prior. This framing underpins spam filters, A/B testing with Bayesian methods, and calibrating a model's confidence.

Where this comes back

  • Week 9Naive Bayes classifiers apply this directly.
  • Week 37Guardrail classifiers on rare attacks face exactly the base-rate problem.

Take many samples of size n from almost any distribution and compute each sample's mean. The Central Limit Theorem says those means form an approximately normal distribution centred on the true mean, with spread σ/n\sigma/\sqrt{n}, the standard error.

That n\sqrt{n} is why precision is expensive: to halve your uncertainty you need four times the data.

It is also why most statistical tests work at all: we can reason about the sampling distribution of a mean even when the raw data is skewed. Watch the histogram of means turn into a bell curve in the simulation, even from a lopsided source.

Xˉn≈N ⁣(μ, σ2n),SE=σn\bar{X}_n \approx \mathcal{N}\!\left(\mu,\ \frac{\sigma^2}{n}\right), \qquad SE = \frac{\sigma}{\sqrt{n}}

Going deeper

The CLT needs finite variance and independent samples. With very heavy tails or strong correlation, convergence is slow or fails, so check the sample size you actually need instead of assuming 30 is enough.

Bootstrapping estimates a sampling distribution empirically: resample your data with replacement thousands of times and recompute the statistic. It's the most practical way to put confidence intervals on eval metrics.

A 95% confidence interval is a procedure that captures the true value 95% of the time across repeated samples. For a proportion pp on nn examples, it is roughly p±1.96p(1−p)/np \pm 1.96\sqrt{p(1-p)/n}.

A hypothesis test sets a null hypothesis H0H_0 (no difference) and asks: if H0H_0 were true, how likely is a result at least this extreme? That probability is the p-value. Small p means the data would be surprising under the null.

What a p-value is not: the probability the null is true, or the probability your result is a fluke, or a measure of effect size. A tiny, useless effect can have p < 0.001 with enough data; a large, important effect can miss p < 0.05 with too little.

CI95%≈p^±1.96p^(1−p^)n\text{CI}_{95\%} \approx \hat{p} \pm 1.96\sqrt{\frac{\hat{p}(1-\hat{p})}{n}}

Going deeper

Statistical power is the probability of detecting an effect that really exists. Underpowered tests (small n, small effects) both miss real improvements and exaggerate the size of the ones they do find. Decide the sample size from the smallest effect you care about.

Multiple comparisons inflate false positives: test 20 prompt variants at p < 0.05 and expect one 'winner' by chance. Correct for it (Bonferroni, Holm) or confirm on a fresh held-out set.

Common pitfalls

  • Running many tests and reporting the one with p < 0.05 (p-hacking).
  • Treating 'not significant' as 'no effect'.

The question you will face constantly: version B scored 3 points higher than A on your eval set. Is that real? With 50 examples, the standard error of an accuracy near 80% is about 0.8⋅0.2/50≈5.7\sqrt{0.8\cdot0.2/50} \approx 5.7 points, so a 3-point gap is well inside the noise.

Better designs: use more examples; use a paired comparison (both versions on the same examples, then look at items where they disagree, as in McNemar's test); and bootstrap confidence intervals by resampling the eval set.

Decide the metric and sample size before looking at results, and keep a held-out set you don't tune against.

Going deeper

For LLM evals, cluster-aware standard errors matter when questions come in groups (several per document), and paired tests (same questions, two systems) remove question-difficulty noise. Anthropic's 'Adding error bars to evals' lays out the recipe.

Peeking at a running A/B test and stopping as soon as it looks significant inflates false positives dramatically. Fix the sample size in advance or use sequential methods designed for peeking.

Best resources for this lesson

Where this comes back

  • Week 36Evals as a deploy gate need exactly this to avoid chasing noise.

NumPy arrays are typed, contiguous blocks of memory. Vectorised operations (a * b, a @ b, np.exp(a)) run in optimised C, often 100× faster than Python loops.

Broadcasting lines up shapes from the right; a dimension of size 1 is stretched to match. So a (1000, 768) matrix minus a (768,) mean vector subtracts the mean from every row, with no loop and no copy.

This is how you compute all pairwise cosine similarities between queries and documents in one line: normalise the rows, then Q @ D.T.

python
import numpy as np
D = np.random.randn(10_000, 384)           # document embeddings
q = np.random.randn(384)                   # query embedding
D /= np.linalg.norm(D, axis=1, keepdims=True)   # broadcasting: (10000,384)/(10000,1)
q /= np.linalg.norm(q)
scores = D @ q                              # cosine similarities, shape (10000,)
top5 = np.argsort(-scores)[:5]

Going deeper

Views vs copies: slicing returns a view that shares memory, so writing to it changes the original; fancy indexing (with arrays or masks) returns a copy. np.shares_memory tells you which you have.

einsum expresses any combination of sums and products in one line (np.einsum('bij,bjk->bik', A, B) is a batched matmul). It's the clearest way to write attention by hand.

Common pitfalls

  • Accidental broadcasting of (n,) against (n,1) produces an (n,n) matrix.

Best resources for this lesson

A DataFrame is a table of typed columns. The workhorses: boolean filtering (df[df.age > 30]), groupby(...).agg(...) for split-apply-combine, merge for SQL-style joins, pivot_table and melt for reshaping, and .isna() for missing data.

EDA is not 'run df.describe()'. Look at distributions (histograms, box plots), relationships (scatter plots, correlation heatmaps), and especially subgroups: a pattern that holds overall may reverse within groups (Simpson's paradox).

A non-obvious finding is something like: one subgroup has a very different distribution, a correlation contradicts intuition, or the way data was collected created an artefact (all missing values come from one source). Those findings drive feature engineering in weeks 7–8.

python
import pandas as pd
df = pd.read_csv("listings.csv")
(df.groupby("neighbourhood")["price"]
   .agg(["count", "median"])
   .query("count > 50")
   .sort_values("median", ascending=False)
   .head(10))

Going deeper

Method chaining (.pipe, .assign, .query) keeps EDA readable and reproducible. Use categorical dtypes and pd.to_datetime early; they make groupbys faster and plots correct.

For data larger than memory, Polars or DuckDB run the same analyses far faster with lazy evaluation, and DuckDB lets you query Parquet and CSV files with SQL directly.

Common pitfalls

  • Chained assignment (df[df.x > 0]['y'] = 1) may silently not modify df; use .loc.

Best resources for this lesson

Where this comes back

  • Week 7Feature engineering starts from what EDA reveals.

Covariance measures whether two variables move together, in their own units. Pearson correlation rescales covariance to [−1, 1], capturing linear association. Spearman correlation uses ranks, catching any monotonic relationship and resisting outliers.

Correlation can be zero for a strong non-linear relationship (y = x² on symmetric data), and high for unrelated variables that share a cause (ice-cream sales and drowning both rise in summer). Always plot the scatter.

Causal claims need an experiment (randomised A/B test) or careful causal reasoning about confounders. In evals: 'longer answers score higher' may mean the judge prefers length, not that the answers are better.

ρX,Y=Cov(X,Y)σXσY=E[(X−μX)(Y−μY)]σXσY\rho_{X,Y} = \frac{\mathrm{Cov}(X,Y)}{\sigma_X\sigma_Y} = \frac{\mathbb{E}[(X-\mu_X)(Y-\mu_Y)]}{\sigma_X\sigma_Y}

Common pitfalls

  • Reading causation into a correlation.
  • Using Pearson on clearly non-linear or outlier-heavy data.
  • Simpson's paradox: a trend that reverses within subgroups.

Where this comes back

  • Week 36LLM-judge biases (like preferring long answers) are correlation traps.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Simulate the CLT

    Draw 10,000 sample means from an exponential distribution for n = 2, 10 and 50. Plot them and compare their spread with σ/√n.

  2. Core

    Bootstrap an eval

    Given 100 per-question scores for two prompts, compute each accuracy with a 95% bootstrap confidence interval and a paired bootstrap CI for the difference. Is the gap real?

  3. Stretch

    Capstone B: EDA

    Full EDA on a messy Kaggle dataset with five non-obvious findings, each backed by a plot. At least one should be a subgroup effect.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 29Monday19 Oct2 h planned

    Probability: distributions, expectation, variance

  2. Day 30Tuesday20 Oct2 h planned

    Bayes theorem and conditional probability

  3. Day 31Wednesday21 Oct2 h planned

    Sampling, CLT, confidence intervals, hypothesis testing

  4. Day 32Thursday22 Oct2 h planned

    NumPy: arrays, broadcasting, vectorised thinking

  5. Day 33Friday23 Oct2 h planned

    Pandas: indexing, groupby, merge, reshape

  6. Day 34Saturday24 Oct3 h planned

    Matplotlib / Seaborn + full EDA on a real dataset

  7. Day 35Sunday25 OctReview

    Review the week, finish anything unfinished, rest

Watch

Statistics FundamentalsPrimary

StatQuest · playlist

Maths & Statistics for ML (Data Analysis Process)

CampusX · playlist

Bayes theorem, the geometry of changing beliefs

3Blue1Brown

But what is the Central Limit Theorem?

3Blue1Brown

p-values: What they are and how to interpret them

StatQuest

Pandas

CampusX · playlist

Pandas Tutorials

Corey Schafer · playlist

Matplotlib Tutorials

Corey Schafer · playlist

Seaborn

CampusX · playlist

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Your new prompt scores 84% vs 81% on 100 eval questions. Ship it?
  • SE of each ≈ 4 points: the gap is within noise
  • Use a paired test on the same questions; look at disagreements
  • Grow the eval set or confirm on held-out data before deciding
Explain a p-value without the common misconception.
  • Probability of data at least this extreme if the null were true
  • Not the probability the null is true
  • Doesn't measure effect size or importance
What is Bayes' theorem and where does it show up in ML?
  • P(H|E) = P(E|H)P(H)/P(E)
  • Base rates matter (rare-event detectors)
  • Naive Bayes, calibration, Bayesian A/B testing

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.A test for a condition with 1-in-1000 prevalence is 99% sensitive and 99% specific. A positive result means the condition is…

  2. 2.To halve the standard error of a mean, you need…

  3. 3.p = 0.03 means…

  4. 4.Prompt B beats prompt A by 3 points on 50 examples (~80% accuracy). Best next step?

  5. 5.Shapes (1000, 768) − (768,) under broadcasting gives…