Peeking in A/B testing: why "already significant" is not a result
Peeking in A/B testing inflates false positives. In 100,000 simulated A/A tests, daily checks turned a 5% error rate into 27.6%. What to do instead.
By Sergey BruhPublished 9 min read
Peeking in A/B testing means checking the results while the test is still running and stopping as soon as the dashboard shows . It is a problem because every extra look gives random noise one more chance to cross the significance line, so you ship far more false winners than the 5% you agreed to. We simulated 100,000 tests in which both versions were identical: one look at the end called 5.0% of them "significant", while a look every day for four weeks called 27.6% "significant" at some point. Opening the dashboard is not the mistake. Making the decision on the first green number is.
A test that turned green on day 9
CatChow, an online pet-food shop, tests a "Subscribe and save" block on the product page. Emily, the product manager, planned the test properly: the baseline conversion is 4%, the smallest lift worth shipping is 10% relative, and that takes about 39,500 visitors per group, or four weeks at 3,000 visitors a day. (The calculation is in our article on A/B test sample size.)
Then the dashboard does this:
On day 9 the variant is 14.4% ahead and the result is significant. Victor, the head of marketing, writes: "It's already green, just ship it." Emily waits. By day 28 the gap has shrunk to 0.9% and .
Here is the catch: this trajectory is one of the 100,000 tests from the simulation below, chosen because it shows the pattern clearly. In every one of them both groups had exactly the same true conversion rate. The block did nothing. Day 9 was noise that looked like a win.
The simulation: 100,000 A/A tests
An A/A test compares a version with itself, so every "significant" difference is a false positive by construction. The settings:
- true conversion rate of 4% in both groups;
- 1,500 visitors per group per day (3,000 a day in total);
- 28 days, so 42,000 visitors per group at the end;
- 100,000 simulated tests, random seed 20261006;
- a two-sided z-test for two proportions with the threshold , the test from the lesson on conversion rate tests.
At every look we calculated the p-value on all the data collected so far and asked one question: was this test "significant" at least once?
One planned look keeps the promise: 5.0%. A harmless-sounding weekly check already gives 12.7%. With daily checks more than one test in four turns green at some point.
Three more facts about the 27,550 false alarms from daily checking:
- They come early. Half of them first crossed the line by day 5, and 60% within the first week.
- They look impressive. At the first crossing the median gap between the groups was 16.8% relative. On a small sample only a large gap can be significant.
- They don't last. Only 18% of these tests were still significant on day 28.
About half of the false alarms favoured B. A team that ships whenever B turns green would have released a do-nothing change in 13.8% of these tests, against 2.5% with a single look.
Why peeking inflates false positives
A p-value is not a progress bar that fills up steadily. While data comes in, the gap between the groups wanders up and down, and the p-value wanders with it, most of all early on, when the sample is small.
The 5% guarantee covers one pre-planned moment: if there is no real effect, the chance that the result sits beyond the line at that moment is 5%. "Stop as soon as it's green" asks a different question: will the result cross the line at any moment? Compare "will it rain at noon next Saturday?" with "will it rain at any time this month?". The second is far more likely, even though the weather is the same.
The rule is also one-sided. A green result ends the test; a grey one gets another try tomorrow. Noise only has to get lucky once.
This is a relative of the multiple comparisons problem, where testing many metrics produces a "winner" by chance. But looks overlap (day 10 contains all of day 9's data), so 28 looks give 27.6% and not the of 28 independent attempts.
What you can monitor during a test
Not looking at all is a bad idea too: a broken variant costs real money.
What you should not do is use the p-value or the lift of the primary metric, the one the decision is based on, as a trigger to stop and ship. For guardrails, agree on the stop conditions before launch ("we stop if checkout errors double"), so that they don't become a second way to peek.
Common mistakes
"We'll ship only if it stays significant two days in a row." In the simulation this rule still flagged 17.2% of A/A tests. Better than 27.6%, but more than three times the nominal rate.
"Not significant yet, let's give it another week." Extending a test until it turns green is the same peeking, just at the other end.
Trusting the lift from a test that was stopped early. Even when the effect is real, the first crossing happens when noise is pushing in your favour, so the lift that goes into the forecast is too optimistic.
What to do in practice
Option 1: fix the sample size and the duration in advance. Calculate the sample with the A/B test sample size calculator, round up to whole weeks and write four things into the test brief: the primary metric, the sample size, the end date and the decision rule. Analyse the primary metric once, at the end, and report the lift together with its confidence interval.
Option 2: use a sequential method. If an early decision is really valuable, use a method designed for repeated looks. The idea is simple: you set the looks in advance and spread the 5% error budget across them, so each look has a stricter threshold than 0.05. In our simulation four weekly looks with at each of them flagged 5.0% of A/A tests, the same as one look at 0.05.
The price is sensitivity. With a real lift from 4.0% to 4.4%, the single look on day 28 found the winner in 82.5% of simulated tests and the four-look scheme in 74.7%, but its tests took 19.7 days on average instead of 28. Many experimentation platforms have a sequential mode built in; check which statistics are behind the green badge in yours.
How to explain it to a stakeholder
Victor doesn't need a lecture on stopping rules. He needs a reason, a date and a promise:
"Green on day 9 is not the result yet. In tests like ours where nothing really changed, more than one in four turns green at some point if you check every day, and four out of five of those are no longer green at the end. We agreed on four weeks because that is what it takes to see a 10% lift. On day 28 I'll bring the final number with its range, and if it's a win, we ship the next morning."
Name the costs too: waiting takes 19 more days, while a false winner means a feature that does nothing and a forecast built on a lift that was never there.
Key takeaways
- Peeking means stopping a test at the first significant result; it turns a 5% false-positive rate into 12.7% with weekly looks and 27.6% with daily looks in our simulation.
- False alarms come early, look large and mostly vanish by the planned end date.
- Monitor bugs, sample ratio and guardrail metrics freely; judge the primary metric only at the planned end.
- Either fix the sample size and duration in advance, or use a sequential method built for repeated looks.
- Agree on the stopping rule before launch, not in front of a green dashboard.
FAQ
Is it OK to look at the dashboard during an A/B test?
Yes, as long as the look can't change the decision. Check that the test runs correctly, that the split is even and that guardrail metrics are healthy. The primary metric is analysed at the planned end, or at the looks your sequential method defines.
Can I stop early if the result is very significant?
Only under a rule set in advance. A far stricter threshold does protect you: daily looks with flagged 0.9% of A/A tests in our simulation. But fixing that threshold before the test is exactly what a sequential method does. Deciding after you have seen the data that the p-value is "small enough" is still peeking.
What if the test is not significant at the planned end?
Stop and report what you have: the observed lift and its confidence interval, which shows how large an effect could still be hiding. Don't extend the test just to chase significance. If the question is still important, plan a new test with a larger sample.
Learn it hands-on
The free lesson Pitfalls: peeking, multiple testing, novelty effect lets you practise on CatChow's banner test that was "already significant" after two days, with quizzes and a practice task checked by AI. It is part of the free course Statistics for Product Managers and Marketers.