1.22 A/B testing and experiment design
You can design an experiment that will actually answer the question.
Before:00. Orientation & SetupUnlocks:03. Data Handling & Analysis04. Classical AI — Agents, Search & Knowledge Representation
A/B testing is how product decisions actually get made, and a good experiment is designed before it starts: sample size fixed, metric chosen, randomisation doing the work of turning comparison into evidence. It follows hypothesis testing because it is that machinery pointed at products. The trap with a name is peeking — checking daily and stopping at the first significant reading inflates false positives dramatically, and sequential methods exist precisely because everyone wants to peek.
Work through these
Randomization, control, blocking
Randomly assigning, holding a comparison group, and grouping similar subjects together before assigning. These three between them remove most of the ways an experiment can mislead.
Sample-size and minimum detectable effect
How many observations are needed to detect an effect of a given size, decided before the experiment starts. Deciding it afterwards means the experiment cannot answer the question.
Peeking, sequential testing, novelty effects
Looking at results early, testing repeatedly, and the temporary excitement any change produces are three ways to get a false positive. Each has a known remedy.
Reading an experiment readout critically
Reading somebody else's experiment critically means asking what was randomised, how many were in it, and what else changed. These questions catch most flawed readouts.
Sign in to keep your progress.
Free resources
We haven't checked most of these for screen reader use yet.
Links last checked 29 Aug 2026.
Stuck here?
Ask a mentor. A real person answers, and they can see exactly which topic you're on. Usually within a couple of working days.
Checking your session…
Topics shown in module order.