Minimum Detectable Effect Calculator
Find out how small a conversion increase your experiment can detect with the traffic and time you have. Free, no account, and nothing you enter leaves your browser.
Power, significance and sample size decide the answer. The free lesson explains them with a worked example. Understand power and sample size
Result
Updates as you type
- Relative uplift
- ≈ +21.67%
- Of the baseline conversion of 5%
- Increase in percentage points
- ≈ +1.08
- From 5% to about 6.08%
- Participants per variant
- 7,000
- 14,000 in the model, A and B together
Other recruitment durations
The same baseline, traffic, alpha and power. Every duration has its own sample and its own result.
7 days
- Relative uplift
- ≈ +31.26%
- Percentage points
- ≈ +1.56
- Target conversion
- ≈ 6.56%
- Per variant
- 3,500
14 daysyour duration
- Relative uplift
- ≈ +21.67%
- Percentage points
- ≈ +1.08
- Target conversion
- ≈ 6.08%
- Per variant
- 7,000
28 days
- Relative uplift
- ≈ +15.11%
- Percentage points
- ≈ +0.76
- Target conversion
- ≈ 5.76%
- Per variant
- 14,000
A longer test detects smaller effects, but it is not always the better plan: it delays the decision and is exposed to more changes in traffic and seasons.
What this estimate assumes
- Conversion
- 5% → about 6.08%
- Available participants
- 14,000 = 1,000 per day × 14 days
- In the model
- 14,000: 7,000 per variant
- Recruitment
- 14 days, without the time participants need to convert
- Significance level
- 5%, two-sided
- Power
- 80%
- Allocation
- Two variants, 50/50
- Method
- Inverse of the sample size formula: normal approximation for two independent proportions, fixed sample size, no continuity correction
How to read this
- With these assumptions, an increase from 5% to about 6.08% corresponds to the selected 80% power in this planning approximation.
- This is not a predicted uplift or a guarantee of a significant result, and 80% is not the probability that B is better.
- A smaller real effect can still be detected sometimes. It just does not reach the requested power under these assumptions.
- MDE describes the sensitivity of the test, not business value. Whether an effect of this size is realistic and worth having is a separate decision.
- Recruitment time does not cover delayed conversions, business cycles such as weekdays and weekends, or changes in traffic.
How the calculator works
The calculator answers the reverse question of a sample size calculator. You already know the traffic and the time you have, and you want to know which increase a test of that size can detect.
It multiplies the daily traffic by the duration, splits the total into two equal groups and searches for the smallest increase for which the sample size formula needs no more participants than one group has. The formula is the normal approximation for two independent proportions with a two-sided test and no continuity correction, the same as in our sample size calculator.
The sample matters less than it seems. Four times the participants only halves the detectable effect: in the example below 7 days give about 31% and 28 days about 15%.
Worked example
A checkout converts 5% of visitors. About 1,000 new participants can enter the experiment each day, and the team has 14 days to recruit them. That is 14,000 participants, 7,000 per variant.
With a 5% significance level and 80% power the minimum detectable effect is about +21.67% relative. That is about +1.08 percentage points: from 5% to about 6.08%.
If the change is unlikely to lift conversion by a fifth, this test will probably miss its effect. The honest options are more time, a page with more traffic, a bolder change, or accepting that this question cannot be answered with an A/B test now.
Assumptions and limits
- One binary outcome per participant: converted or not. Revenue and other continuous metrics need a different method.
- Participants are randomized independently and counted once. Sessions, page views and repeat visits are not participants.
- Every participant is observed for the same length of time after entering the test.
- The analysis is planned in advance and done once, at the end. Repeated peeking, several metrics and several comparisons need corrections that are not included.
- Two variants with a 50/50 split. With three or more variants the detectable effect is larger than shown here.
- Increases only. Decreases and metrics where lower is better are not covered.
- The same average traffic every day. Real traffic varies, so the real sample may be smaller or larger.
- The result comes from an approximation. With rare conversions, conversions close to 100% or few participants it does not hold, and the calculator shows no number.
Questions and answers
- Is MDE the uplift my change will produce?
- No. MDE says how sensitive the experiment is, not what the change will do. It is the smallest real effect the test detects with the chosen power. The actual effect may be larger, smaller or absent.
- What is the difference between relative uplift and percentage points?
- Relative uplift is a share of the baseline: +20% from 5% gives 6%. Percentage points are added to the baseline: +1 point from 5% also gives 6%, while +20 points would give 25%. The calculator always shows both, together with the baseline and the target.
- Is the traffic per variant or for both variants?
- For both variants together. Enter all new participants who enter the experiment per day, and the calculator splits them 50/50. If only a part of the visitors is eligible or included in the test, enter that part.
- What do alpha and power change?
- Alpha is the accepted risk of a false alarm, power is the chance of detecting a real effect of the MDE size. A stricter alpha or a higher power makes the detectable effect larger for the same sample. Choose both before the test, not after seeing the result.
- Is the duration the whole length of the experiment?
- It is the recruitment window: the days when new participants enter. If a conversion takes time, for example a purchase a week after the first visit, the experiment lasts longer than the recruitment. Whole weeks are usually a good idea, because weekdays and weekends differ.
- Why is there no result for a rare or a very high conversion?
- The method is an approximation that needs enough expected conversions and non-conversions in each variant, at least 10 of each. With fewer the number would look precise and mean little, so none is shown. That does not make an experiment impossible: it needs more participants or an exact method.
- Why do other calculators show a different MDE?
- Calculators differ in the formula, one-sided or two-sided tests, continuity corrections, and in whether they solve the power equation or invert a sample size formula, as this one does. Compare results only when the settings match. A difference of a few percent is normal.
The reverse question
Do you already know the effect you need to detect? Calculate how many participants that takes. A/B test sample size calculator