Skip to content
Log in
← Statistics glossary

holdout group

Also called: holdout, long-term holdout, holdout test, control holdout, global holdout

A small group kept without a feature or campaign for a long time to measure its cumulative effect after launch.

A holdout group is a randomly chosen slice of users, usually 1–10%, who don't get a feature or a campaign while everyone else does, and who stay that way for weeks or months. Comparing the rest with the holdout shows the long-term, cumulative effect of what you shipped.

It answers questions a normal A/B test can't:

  • Does the effect last? A two-week test catches novelty; a holdout shows whether the lift survives after three months.
  • What did a whole set of changes add up to? A global holdout kept away from every launch of a quarter shows whether ten "winning" tests really added up to ten wins.
  • Is a campaign incremental? In marketing, a holdout that receives no emails, pushes or ads shows how many conversions would have happened anyway. This is the cleanest way to measure incrementality.

The effect is read like any experiment: the difference between exposed users and the holdout, with its confidence interval:

Δ=pexposed−pholdout,SE=pe(1−pe)ne+ph(1−ph)nh\Delta = p_{\text{exposed}} - p_{\text{holdout}}, \qquad SE = \sqrt{\frac{p_e(1-p_e)}{n_e} + \frac{p_h(1-p_h)}{n_h}}

The price is real: holdout users miss something that works, so keep the group as small as the needed precision allows and set an end date up front. Randomize at the right level too. If a feature affects everyone in a shared space, like a group of friends, hold out whole groups, or the effect leaks into the holdout through their friends.

Example

Halves launches automatic settle-up reminders to 40,000 active users but holds out 5% of groups (2,000 users), so no one in a holdout group gets reminders from a friend's side either. After 90 days, the metric is "recorded a settle-up in the last 30 days":

Group
Users
Settled up
Rate
Reminders
38,000
13,680
13,680 / 38,000 = 36.0%
Holdout
2,000
640
640 / 2,000 = 32.0%
  1. Effect: Δ=36.0%−32.0%=4.0\Delta = 36.0\% - 32.0\% = 4.0 pp (relative lift 4.0/32.0=12.5%4.0 / 32.0 = 12.5\%).
  2. Standard error: SE=0.36×0.6438000+0.32×0.682000=0.00000606+0.0001088≈0.0107SE = \sqrt{\frac{0.36 \times 0.64}{38000} + \frac{0.32 \times 0.68}{2000}} = \sqrt{0.00000606 + 0.0001088} \approx 0.0107.
  3. 95% interval: 4.0±1.96×1.07≈4.0±2.14.0 \pm 1.96 \times 1.07 \approx 4.0 \pm 2.1, i.e. +1.9 to +6.1 pp.

The launch test showed +6 pp in its first two weeks; three months later the effect is smaller but clearly above zero. The cost of the holdout was about 2000×0.04=802000 \times 0.04 = 80 users who didn't settle up in a month, which Maya judged worth a trustworthy long-term number. (With users grouped, a precise interval should account for clustering, which widens it somewhat.)

Common mistakes

  • A holdout that isn't random. "Users on old app versions" or "one country" differ in more than the feature; pick the holdout by random assignment.
  • Leaks into the holdout. Friends in the same group, shared devices or a marketing email forwarded to a colleague expose holdout users; randomize at the level where exposure spreads.
  • Too small to measure anything. A 1% holdout of a small product can't detect a realistic effect; size it for the effect you care about.
  • Letting it run forever. Holding people out of a working feature has a cost; set the end date and the metric before launch.
  • Reading many metrics until one moves. Name one primary metric up front; the rest are guardrails.