Skip to content
Log in
← Statistics glossary

cohort analysis (cohort table)

Also called: cohort table, cohort analysis, cohort matrix, retention table, cohort retention table, retention cohort analysis

A table with one row per acquisition cohort and one column per period since joining, showing retention or revenue. Read rows for a cohort's life, columns to compare cohorts.

Cohort analysis groups users by when they started (sign-up week, first purchase month) and follows each group separately over time. Its standard view is the cohort table: one row per cohort, one column per period since joining (week 0, week 1, week 2…), and in each cell the share of the cohort that was active in that period, or the revenue it brought. The table is triangular, because newer cohorts haven't lived through as many periods yet.

It lets you read three directions:

  • Along a row: how one cohort decays over time and whether its curve flattens.
  • Down a column: whether newer cohorts retain better than older ones at the same age. This is how you see whether product changes are working.
  • Along a diagonal: cells that fall in the same calendar week. If a whole diagonal dips, something happened on that date (an outage, a holiday) rather than in the product's long-term quality.

Averaging a column needs care. Weight each cohort by its size, and include only cohorts that have reached that period:

rˉk=∑cactivec,k∑csizecover cohorts c old enough to have period k\bar r_k = \frac{\sum_c \text{active}_{c,k}}{\sum_c \text{size}_c} \quad \text{over cohorts } c \text{ old enough to have period } k

A simple mean of percentages lets a small cohort count as much as a large one. Dividing by every cohort's size, including ones too young to have the period, drags the average down (this is right-censoring). Beyond the average, cohort tables are the basis for comparing segments (channel, platform, group type), spotting a retention plateau and forecasting active users.

Example

Halves weekly cohorts, share of each cohort active in week k after sign-up, measured on July 6:

Cohort (sign-up week)
Users
Week 1
Week 2
Week 3
Week 4
Jun 1
1,000
420 (42%)
330 (33%)
290 (29%)
270 (27%)
Jun 8
1,500
600 (40%)
450 (30%)
390 (26%)
—
Jun 15
500
240 (48%)
190 (38%)
—
—
Jun 22
800
344 (43%)
—
—
—

Average week-2 retention, three ways:

  1. Weighted, eligible cohorts only (correct): 330+450+1901000+1500+500=9703000=32.3%\frac{330 + 450 + 190}{1000 + 1500 + 500} = \frac{970}{3000} = 32.3\%.
  2. Simple mean of percentages: 33+30+383=33.7%\frac{33 + 30 + 38}{3} = 33.7\%. The small June 15 cohort pulls it up.
  3. Dividing by all cohorts, including June 22: 9703800=25.5%\frac{970}{3800} = 25.5\%. The June 22 users haven't reached week 2 yet, so this understates retention by almost 7 pp.

Reading down the week-1 column, the June 15 cohort (48%) stands out. Maya checks its source before celebrating: it came mostly from group invites, which retain better than store traffic.

Common mistakes

  • Averaging percentages instead of users. Weight by cohort size, or a tiny cohort sways the result.
  • Including cohorts that haven't reached the period. Empty cells are "not yet", not zero.
  • Cohorts too small to read. With 50 users per cohort, a 6 pp swing is noise; use wider cohorts (weeks, months) or show intervals.
  • Ignoring the mix behind a cohort. A cohort from a partner promo or a summer trip season differs in who joined, not in how good the product became.
  • Mixing definitions of "active" between tables. Compare only tables built with the same activity event, period and time zone.