Don't peek
Run an A/B test where nothing changed. Check it every day and watch fake winners appear.
Checked once, day 20
–
false winners
Peeked every day
–
false winners
With no real difference, a fair test should cry wolf about 5% of the time. Run 1,000 tests and compare the two columns.
What’s going on
Each test compares two conversion rates with a two-proportion z-test: z = (p̂B − p̂A) / √[p̂(1 − p̂)(1/nA + 1/nB)], where p̂ pools both groups. A p-value under 0.05 is called significant.
That 5% promise holds for one look at a sample size you fixed in advance. Check every day and stop the first time p dips under 0.05, and you give chance twenty tries to fool you. With no real difference, the simulation shows the false winner rate jumping from about 5% to roughly a quarter of all tests.
Fixes used in practice: decide the sample size before you start and look once, or use a sequential test built for repeated looks. Either way, a significant result you went fishing for is weaker evidence than it looks.