AI Engineer Path

Concept Lab · Week 5 · Prerequisites

Is prompt B really better?

Simulate eval sets of different sizes and see how often noise fools you.

The idea: A/B testing for model and prompt changes

The question you will face constantly: version B scored 3 points higher than A on your eval set. Is that real? With 50 examples, the standard error of an accuracy near 80% is about 0.8⋅0.2/50≈5.7\sqrt{0.8\cdot0.2/50} \approx 5.7 points, so a 3-point gap is well inside the noise.

Better designs: use more examples; use a paired comparison (both versions on the same examples, then look at items where they disagree, as in McNemar's test); and bootstrap confidence intervals by resampling the eval set.

Decide the metric and sample size before looking at results, and keep a held-out set you don't tune against.

Open the full lesson in week 5

Next simulation: Bias–variance tradeoff