Skip to content
Log in

How long to run an A/B test: from sample size to whole weeks

How long to run an A/B test: divide the required sample by eligible daily traffic, round up to whole weeks, and plan for novelty effects and sales.

By Sergey BruhPublished 9 min read

To work out how long to run an A/B test, divide the total sample the test needs by the number of eligible users who enter it each day, then round the result up to whole weeks. Run it for at least two full weeks even when the maths says four days, and treat anything beyond about eight weeks as a sign that the test needs a different design.

"Can we have the result by Friday?"

CatChow, an online pet-food shop, wants to show a delivery estimate ("Arrives on Thursday") next to the buy button on product pages. Emily, the product manager, plans the test:

  • baseline: 5% of visitors who open a product page place an order;
  • minimum detectable effect (MDE): a 10% relative lift, from 5.0% to 5.5%;
  • significance level 5% (two-sided), power 80%.

The sample size calculator returns 31,234 users per group, 62,468 in total. The post on A/B test sample size explains where that number comes from.

The shop gets about 8,000 visitors a day, but only 2,500 of them open a product page. So:

days=required sampleeligible users per day=62,4682,500≈25\text{days} = \frac{\text{required sample}}{\text{eligible users per day}} = \frac{62{,}468}{2{,}500} \approx 25

Rounded up to whole weeks, the test runs for 4 weeks (28 days). Friday is not happening.

Two rules behind the number

Count only eligible traffic

Eligible users are those who reach the tested page and can see the change. Had Emily divided by all 8,000 visitors, the plan would say 8 days. In 8 days only 20,000 people open a product page, 10,000 per group, and with that sample the chance of detecting a real 10% lift is about 35% instead of the planned 80%. The test would most likely come back "not significant", and a working idea would be shelved.

Take the traffic figure from ordinary recent weeks and count unique users, not sessions.

Round up to whole weeks

People shop differently on a Tuesday morning and on a Sunday evening. If a test launched on a Monday stopped on day 25, Monday to Thursday would be counted four times and Friday to Sunday only three. Whole weeks give every weekday the same weight.

From traffic and MDE to weeks

The same test (baseline 5%, α=0.05\alpha = 0.05, power 80%) for different effects and traffic levels:

Relative MDE
Users per group
1,000 eligible users/day
2,500/day
5,000/day
5%
122,124
245 days → 35 weeks
98 days → 14 weeks
49 days → 7 weeks
10%
31,234
63 days → 9 weeks
25 days → 4 weeks
13 days → 2 weeks
15%
14,193
29 days → 5 weeks
12 days → 2 weeks
6 days → 1 week
20%
8,158
17 days → 3 weeks
7 days → 1 week
4 days → 1 week

Doubling the traffic halves the duration, but halving the MDE roughly quadruples it.

If you start from a deadline, turn the question around. The MDE calculator shows that with 2,500 users a day CatChow can detect a lift of 13.5% in two weeks and 9.4% in four. Which effect is worth chasing is the topic of the post on the minimum detectable effect.

The shortest and the longest sensible test

Minimum: one to two full business cycles. For most online shops the cycle is a week, so "1 week" in the table becomes two in practice. A single week can be an odd one (a holiday, a newsletter, an outage), and two weeks let you compare the first with the second. If your customers buy on a longer rhythm, such as a monthly order after payday, cover one full cycle of it.

Maximum: about six to eight weeks. Longer tests degrade:

  • Cookie loss and new devices. A returning person with cleared cookies or a new phone is randomised again and may land in the other group. The more people see both versions, the more the difference blurs.
  • A shifting audience. Frequent visitors enter the test in its first days; later arrivals are increasingly occasional visitors who behave differently.
  • Seasonality and other launches. Every extra week adds a chance that a price change, a campaign or a competitor's sale lands inside your test.

If the plan says 14 weeks, don't wait it out. Accept a larger MDE, test a bolder change, or move to a metric with a higher baseline (add-to-cart instead of order).

Novelty and primacy effects

A novelty effect is a lift that exists because the change is new: regular visitors notice the unfamiliar block, click it, and later lose interest. A primacy effect is the opposite: people used to the old version stumble at first, and the variant improves as they adapt.

To spot them, compare the first week with the later ones. Suppose Emily's test finished like this, with 8,750 users per group each week:

Week
Control
Variant
Relative lift
1
4.98%
6.40%
+28.4%
2
5.04%
5.55%
+10.2%
3
4.95%
5.42%
+9.5%
4
5.03%
5.39%
+7.3%
All 4 weeks
5.00%
5.69%
+13.8%

Weekly numbers are noisy: each weekly difference carries a margin of about ±0.7 percentage points, so the wobble between weeks 2 and 4 means little. A first week with a lift almost four times the last one's is a different matter. Weeks 2 to 4 together give +9.0%, a more honest estimate of what CatChow will keep.

A second check: split new and returning users. Newcomers have no old version to compare with, so a lift seen only among returning users is probably novelty.

Traffic allocation: 50/50 or 90/10?

Showing the variant to only 10% of users feels safer, but it is expensive. For the same test (5.0% to 5.5%, 2,500 eligible users a day):

Split (control / variant)
Total users needed
Days
Rounded up
50 / 50
62,468
25
4 weeks
80 / 20
96,535
39
6 weeks
90 / 10
170,970
69
10 weeks

The calculator assumes an even split; the other two rows apply the same formula to unequal groups. The smaller group limits the precision of the comparison: at 90/10 the variant still needs about 17,100 users, and at 250 a day that takes ten weeks. In total, a 90/10 split costs 2.7 times the traffic of an even one.

If the risk worries you, expose 10% of users for a day or two to catch bugs, then switch to 50/50 and count the test from that moment.

Sales, holidays and campaigns

A test measures the effect on the people who were there while it ran. Black Friday visitors hunt for discounts and buy with a different urgency, so a lift measured that week may not hold in February.

  • Don't start or finish a test inside a sale unless the sale itself is what you are testing.
  • Check the marketing calendar: a large email or ad campaign brings a burst of unusual traffic.
  • If a campaign lands inside the test anyway, finish as planned and report the result with and without those days.

Stopping early: only for harm

Stopping the moment the primary metric looks significant is peeking, and it inflates false positives; the post on A/B test peeking explains why. Stopping because the variant hurts users is different, as long as the rule is written down before launch:

  • Choose two or three guardrail metrics that must not get worse: payment errors, page load time, cancellations, support contacts.
  • Set a threshold for each and make it strict, for example p<0.001p < 0.001 instead of 0.05, because you will check it every day.
  • Stop at once for a plain bug, such as a broken button in one browser, then fix it and restart with fresh data.

This is safe because an early stop for harm can only end in "don't ship", never in shipping a false winner.

Common mistakes

  • Dividing by total site traffic instead of the users who reach the tested page.
  • Stopping on the day the sample is reached, in the middle of a week.
  • Adding "a few more days" because the result is almost significant. That is peeking too.
  • Judging by the first days, when novelty is strongest.

Pre-launch checklist

  1. The primary metric, baseline and MDE are written down.
  2. The sample is calculated for the split you will actually use.
  3. Daily traffic is counted in eligible unique users, from ordinary weeks.
  4. The duration is in whole weeks: at least two, at most about eight.
  5. The calendar is clear of sales, holidays and big campaigns.
  6. Guardrail metrics and stop rules for harm are agreed.
  7. The end date is fixed.

Key takeaways

  • Duration = required sample / eligible users per day, rounded up to whole weeks.
  • Run for at least two weeks, and redesign any test that needs more than about eight.
  • Compare the first week with the later ones to catch novelty and primacy effects.
  • An uneven split is costly: 90/10 needs 2.7 times the traffic of 50/50.
  • Stop early only for harm, by rules set before launch.

FAQ

Can I run an A/B test for less than a week?

Only to catch bugs. Three days cover only part of the weekly pattern, and the first days carry the strongest novelty effect. Run one full week at the very least, and two whenever you can.

What if the calculation says the test will take three months?

Change the test, not the calendar: a larger MDE with a bolder change, a metric higher in the funnel, or a page with more traffic. If none of that works, decide without a test and say so openly.

Does the test end when the last user enters it?

Not quite. People who joined on the last day need as much time to decide as those who joined on the first. If most orders arrive within three days of the first visit, stop adding new users on day 28 and keep counting orders for three more days.

Learn it hands-on

The free lesson Pitfalls: peeking, multiple testing, novelty effect works through peeking, multiple comparisons and the novelty effect on CatChow experiments, with quizzes and a practice task checked by AI. It is part of the free course Statistics for Product Managers and Marketers.

Learn it in the course

Learn statistics hands-on

A free course for PMs and marketers: short lessons, real product data and exercises with instant feedback.

Start the free course

Related articles