Concept Lab · Week 5 · Prerequisites
Is prompt B really better?
Simulate eval sets of different sizes and see how often noise fools you.
The idea: A/B testing for model and prompt changes
The question you will face constantly: version B scored 3 points higher than A on your eval set. Is that real? With 50 examples, the standard error of an accuracy near 80% is about points, so a 3-point gap is well inside the noise.
Better designs: use more examples; use a paired comparison (both versions on the same examples, then look at items where they disagree, as in McNemar's test); and bootstrap confidence intervals by resampling the eval set.
Decide the metric and sample size before looking at results, and keep a held-out set you don't tune against.
Next simulation: Bias–variance tradeoff