Pitfalls: peeking, multiple testing, novelty effect
The most common ways real A/B tests go wrong, and how to avoid them.
Lesson 12 of 24~27 min of learningIncludes ~20 min for questions and tasks
Contents1 of 48 steps
You
Alex Tutor
You
Alex Tutor
Alex Tutor
Alex Tutor
Question 1
Stopping a test as soon as it first reaches p < 0.05, then declaring victory, keeps your actual false-positive rate at 5%.
You
Alex Tutor
Alex Tutor
Question 2
You test 5 independent guardrail metrics, each at α = 0.05, and none of them truly changed. What is the probability that at least one shows a false positive, as a percentage?
Units: % (e.g. 10.0)
Alex Tutor
You
Alex Tutor
Alex Tutor
Question 3
Using the Bonferroni correction for 5 comparisons at a family-wise α = 0.05, what per-test significance threshold should you use?
Units: alpha (e.g. 0.01)
Alex Tutor
You
Alex Tutor
Alex Tutor
Question 4
What's the safest way to handle wanting to check results before your pre-computed sample size is reached?
Alex Tutor
You
Alex Tutor
You
Alex Tutor
Question 5
Match each pitfall to its description.
Tap an answer, then tap the row it belongs to. You can also drag.
- Peeking
- Multiple comparisons problem
- Novelty effect
- Sample ratio mismatch (SRM)
Answers left to place
You
Alex Tutor
You
Alex Tutor
Question 6
Which of these are classic causes of a false 'winning' variant? Select all that apply.
Alex Tutor
You
Alex Tutor
You
Alex Tutor
Question 7Short answer · AI-checked task
CatChow's dashboard shows the banner test 'already significant' after 2 days, but the pre-registered sample size needs 2 weeks. In 2-3 sentences, explain to Sofia's teammate why they shouldn't ship yet.
Write your answer and get a score with feedback from our AI reviewer.
Log in to get AI feedbackYou
Alex Tutor
You
Alex Tutor
Question 8Case · AI-checked task
Write a one-paragraph recommendation for the team.
CatChow's 'gift-wrap' upsell test ran for 3 weeks, tracking 12 metrics without any correction for multiple comparisons. Day 1 showed a huge lift in the primary metric, which shrank to roughly zero by week 3. Separately, one of the 12 metrics - email sign-ups - came out significant at p = 0.03; none of the others did. Write a one-paragraph recommendation for the team.
Write your answer and get a score with feedback from our AI reviewer.
Log in to get AI feedbackYou
Alex Tutor
Alex Tutor
That’s the lesson. You answered every task — nicely done.