A random variable maps outcomes to numbers. Discrete ones have a probability mass function (PMF); continuous ones a probability density function (PDF), where only intervals have non-zero probability: , but is the area under the curve.
The expectation is the probability-weighted average. Variance is the expected squared distance from the mean; its square root is the standard deviation. Squaring penalises large deviations heavily, for exactly the same reason mean squared error does.
Distributions to know: Bernoulli (one yes/no trial: is this answer correct?), Binomial (count of successes in n trials: correct answers on an eval set), Normal (sums and averages of many things, know it cold), Uniform, and Poisson for counts per interval.
Going deeper
The variance of a sum of independent variables is the sum of their variances, which is why averaging n samples shrinks variance by 1/n. Correlated samples (for example, several eval questions generated from the same document) shrink it much less.
Heavy-tailed distributions (latencies, costs, token counts) make the mean misleading. Report medians and high percentiles (p95, p99) for anything operational.
Best resources for this lesson
- InteractiveSeeing Theory: probability distributions · Brown University's visual introduction
- CourseKhan Academy: random variables
Where this comes back
- Week 36Eval accuracy is a binomial proportion with a standard error you can compute.