Cohort Retention Calculator
Paste your cohort table and see the heatmap, the average curve and where it flattens.
A retention table is a chart before it is a number. Learn to read charts without fooling yourself
Learn to read charts without fooling yourself
Your data stays in this browser. Clear it with Start empty.
How to read a cohort table
Each row is a cohort: the users who started in the same day, week or month. Each column is the time since they started: week 1, week 2 and so on. A cell says what share of the cohort was active in that period.
Read across a row to see how one cohort fades. Read down a column to compare cohorts at the same age: if the newer rows are darker, the product keeps people better than it used to. The diagonals are calendar time, so a dark or pale diagonal points to something that happened on a date (an outage, a campaign, a holiday) rather than to a change in the product.
The newest cohorts have the fewest columns because they have not lived long enough. Those cells are hatched: they have no data yet, and they are not zeros.
Why the average is weighted and the tail is dashed
A plain average of the percentages lets a cohort of 50 users count as much as a cohort of 5,000. The default average adds up the active users and divides by the users who started, over the cohorts that reached the period. Switch to "Each cohort counts equally" when every cohort matters the same to you, for example when cohorts are markets.
Periods a cohort has not reached are left out, not counted as zeros. Counting them as zeros would make the curve drop at the end for no reason.
Late periods rest on the oldest cohorts only. When fewer than three cohorts back a period, the curve is dashed and the table says how many cohorts it is based on. Read those points as a hint, not as the answer.
Method
- Cell retention = active users in the period ÷ users who started in the cohort.
- Weighted average of a period = Σ active users ÷ Σ cohort sizes, over the cohorts that have a value for the period.
- Simple average of a period = the mean of those cohorts' percentages.
- Flattening: the first period after which the weighted curve changes by less than 1 percentage point per period for the next two periods, using only periods backed by at least 3 cohorts. If the periods that could show it rest on fewer cohorts, the tool says it is too early.
- Newer vs older: the cohorts that reached the chosen period, in table order, split into an older and a newer half (the middle one left out when their number is odd), each half weighted by size. Within 1 percentage point reads as about the same.
- Colour: six steps from 0% to the highest retention in the table. The number is always printed in the cell.
Questions and answers
- Should I enter counts or percentages?
- Counts, if you have them: the averages then use exactly your numbers. Percentages work too; the cohort size is still needed so that bigger cohorts weigh more. Switching the input converts the table, and Undo brings it back.
- My tool shows rolling retention. Does that work?
- Yes, the calculator reads whatever you paste. Classic retention counts users active in exactly that period; rolling (or unbounded) retention counts users active in that period or any later one, so it never goes up and is always higher. Label your data so readers know which one they see, and don't compare one kind with the other.
- Does a flat curve mean product–market fit?
- A curve that levels off means some users keep coming back, which is a good sign. How high the plateau should be depends on the product, how often people need it and how you count activity; there is no universal threshold. Compare with your own earlier cohorts first.
- Why is there no significance test for newer vs older cohorts?
- Cohorts differ in season, channel mix and size, and the same users are counted in several periods, so a simple test would give false confidence. The tool describes the difference; to prove that a change caused it, run an experiment.
- What format does the paste need?
- One row per cohort, oldest first: the cohort name, the number of users who started, then one column per period. Leave cells empty where the cohort has no data yet. Spaces in numbers (1 234) are fine, and so is 1,234 in English when the columns are separated by tabs or semicolons.
- Where is my data stored?
- Only in this browser, so the table is still here after a reload. Nothing is sent to us or anyone else. Start empty clears it.
Related tools
- TAM, SAM and SOM worksheet
Estimating market size bottom-up with a realistic range.
- Confidence interval calculator
Seeing how precise a single conversion rate is.