Unit Economics Wiki

Cohort Analysis for Founders

Cohort analysis follows customers who share a start event and measures them at the same age, so growth is never confused with retention.

Unit EconomicsUpdated Aug 13, 202611 min read

Snapshot

What it is

Cohort analysis groups customers by a shared start event — signup, activation, first paid invoice, contract start — and measures each group at the same age rather than on the same calendar date.

Why it matters

New sales can hide churn in older cohorts indefinitely. A cohort grid is the only view that separates "we are acquiring more" from "we are keeping more," and it is the denominator that makes CAC, payback, and LTV auditable rather than rhetorical.

What it is not

A calendar trend report. Charting monthly active users or monthly revenue over time is not a cohort view, because every point mixes customers of different ages.

Key takeaways

  • A defensible cohort states four things: who enters, what starts the clock, what outcome is measured, and how incomplete observation is handled.

  • Logo retention, GRR, and NRR answer different questions. GRR cannot exceed 100%; NRR can. Never report one and call it the other.

  • Attach economics to the cohort that generated them. Acquisition cost belongs to the cohort it bought; recovery is measured in gross profit, never revenue.

  • Censoring is not churn. Accounts still active at the data cutoff must be removed from the denominator, not counted as losses.

  • Report counts alongside percentages. A 50% rate on eight accounts is not evidence.

What is a cohort?#

A cohort is a fixed membership set created by a rule. A January paid cohort might include every account whose first recurring invoice started in January. "Month 3" then means the third completed month after each account's own start date, not March for everyone.

Three measures do most of the work:

MeasureFormulaWhat it answersCeiling
Logo retention at age t
active original members at t ÷ original members
Are we keeping customers?
100%
Gross revenue retention (GRR)
(starting recurring revenue − churn − contraction) ÷ starting recurring revenue
Are we keeping revenue, ignoring upsell?
100%
Net revenue retention (NRR)
(starting recurring revenue − churn − contraction + expansion) ÷ starting recurring revenue
Is the installed base growing on its own?
None

The gap between GRR and NRR is the expansion story. The gap between logo retention and GRR is the account-size story: losing many small accounts and keeping a few large ones produces weak logo retention and strong GRR.

Definitions must specify whether reactivations, usage overage, foreign exchange, acquired customers, and one-time services are included. Public companies disclose these choices precisely because the choices move the number. Vertex, for example, reports NRR of 105% and GRR of 94% for the same customer base at the same date — one number above 100% and one below, describing the same year. (Vertex Q4 2025 results)

The exact policy matters less than having a stable, versioned, auditable one. Never silently change the definition when the chart becomes inconvenient.

Why does cohort analysis matter to founders?#

It separates acquisition from product quality. Total active customers can rise every month while each successive cohort retains worse. Only an age-aligned grid shows whether growth comes from more entrants or better survival.

It makes go-to-market changes testable. Channel, segment, contract type, implementation path, and even the closing rep produce different retention and payback. A cheaper lead source that produces customers who churn in month four is not cheaper.

It exposes changing unit economics before the P&L does. Acquisition cost belongs to the cohort that generated it. Blending an old, efficient cohort with a new, expensive motion conceals the deterioration for several quarters.

It gives fundraising claims a denominator. A single company-wide retention number cannot show trajectory. Investors want to see whether recent cohorts activate faster, retain better, and expand more than older ones — and whether the recent cohorts are old enough to say.

How do you build a cohort analysis?#

1. Write the cohort contract#

Before any SQL, write down: the entity (user, account, workspace, location, contract); the entry event and timestamp; eligibility and exclusions; the period grain (day, week, month, quarter); the outcome (activity, logos, recurring revenue, gross profit); the treatment of upgrades, downgrades, pauses, reactivations, and mergers; and the data-maturity lag. Version it.

2. Build an age-based matrix#

Rows are entry periods; columns are age since entry. Show the eligible denominator and the absolute count next to every percentage.

3. Plot survival, handling censoring properly#

Naive retention rates work only when every account has been observed for the full window. When accounts enter at different times, or some are still active at the data cutoff, those accounts are censored — you know they survived at least this long and nothing more. Treating them as churned understates retention; leaving them in the denominator overstates it.

The Kaplan–Meier product-limit estimator handles this by re-computing the at-risk population at each event time:

S(t) = Π over event times j ≤ t of (1 − d_j / n_j)

where d_j is the number of churn events at time j and n_j is the population still at risk immediately before it. Kaplan and Meier formalised exactly this treatment of incomplete observations in 1958, and it remains the correct default for young cohorts. (Kaplan & Meier, 1958)

4. Attach economics#

cohort payback month = first month in which
    cumulative cohort gross profit ≥ cohort acquisition cost

Use gross profit, not revenue — revenue does not repay acquisition cost when delivery has a cost. Include onboarding and channel cost consistently. If payback has not occurred inside the observed window, report that fact rather than extrapolating a straight line through the censoring boundary.

5. Segment only after the base view works#

Useful cuts: segment, use case, channel, geography, plan, implementation type, contract length, campaign. Stop slicing while the cells still contain enough accounts to mean anything.

Key Facts

01

NRR and GRR can bracket 100% for the same customers in the same period

Vertex reported NRR of 105% and GRR of 94% at 31 December 2025, alongside ARR of $671.0 million and average annual revenue per direct customer of $137,867. Reporting only the 105% would have hidden a 6-point annual revenue loss before expansion.

Vertex Q4 2025 results
02

Segment cohorts diverge sharply from the blended average

Similarweb reported overall NRR of 98% in Q4 2025 while NRR for customers with ARR of $100,000 or more was 103% — and that group, 454 customers out of 6,128, generated 63% of total ARR. The blended number describes almost none of the revenue.

Similarweb Q4 and FY2025 results
03

Censoring has a standard, 68-year-old solution

Kaplan and Meier's product-limit estimator, published in the Journal of the American Statistical Association in 1958, estimates survival without assuming a functional form and without treating unobserved futures as failures.

Kaplan & Meier, 1958
04

Regulators expect a written metric definition

SEC interpretive release 33-10751 (issued 30 January 2020) states that a company disclosing a performance metric should provide a clear definition of the metric and how it is calculated, why it is useful to investors, how management uses it, and any change in the methodology. That is a reasonable bar for a board deck too.

SEC Release 33-10751

Worked example#

Northwind's January cohort begins with 200 customers and $20,000 of monthly recurring revenue.

Retention at month 6. 150 customers remain. Churned customers represented $5,000 of starting MRR; survivors contracted by $1,000 and expanded by $4,000.

logo retention = 150 / 200                              = 75.0%
GRR  = ($20,000 − $5,000 − $1,000) / $20,000  = $14,000 / $20,000 = 70.0%
NRR  = ($20,000 − $5,000 − $1,000 + $4,000) / $20,000 = $18,000 / $20,000 = 90.0%

The cohort lost a quarter of its logos and 30% of its starting revenue; expansion among survivors clawed back 20 points. Reporting only "90% retention" would be both ambiguous and flattering.

Payback on the same cohort. The cohort cost $120,000 to acquire and onboard. Monthly gross profit attributable to it ran $15,000, $16,000, $15,500, $14,800, $14,000, $13,500, $13,000, and $12,500 in months 1–8.

cumulative gross profit through month 8 = $114,300   → below $120,000, no payback yet
month 9 contributes $12,000 → cumulative $126,300     → payback occurs in month 9

Note what would have happened using revenue instead of gross profit: at an 80% gross margin the cohort would have appeared to cross $120,000 in month 7, two months early. The margin boundary is not a rounding detail.

Survival with censoring. Twenty of the 200 customers churn in month 1:

S(1) = 1 − 20/200 = 90.0%

Five accounts become unobservable before month 2 (data cutoff, migration, acquisition), leaving 175 at risk; 15 churn:

S(2) = 90.0% × (1 − 15/175) = 90.0% × 0.9143 = 82.3%

Ten more are censored before month 3, leaving 150 at risk; 12 churn:

S(3) = 82.3% × (1 − 12/150) = 82.3% × 0.92 = 75.7%

The 15 censored accounts are counted neither as retained forever nor as immediate churn. Had they been treated as churn, month-3 survival would have read 71.5% instead of 75.7% — a four-point error introduced entirely by a data-window artefact.

A caveat on precision. All three retention figures above are point estimates from a 200-account cohort. A cohort of 120 accounts with 7 churn events by month 6 gives a point survival estimate of 94.2%, but a 95% Wilson interval of roughly 88.5% to 97.2%. Report the interval, or at least the counts, before anyone builds a five-year forecast on the midpoint.

What are the common mistakes?#

MistakeWhy it breaksFix
Grouping by calendar activity rather than entry event
Produces a trend line, not a cohort; ages are mixed
Anchor every row to the entry event
Retroactively changing cohort membership
The chart stops being reproducible
Version the definition; restate explicitly
Calling NRR "retention"
Expansion masks churn; NRR > 100% with 60% logo retention is common
Always publish logo retention, GRR, and NRR together
Including one-time services in recurring retention
Implementation revenue lands as fake "expansion"
Exclude non-recurring lines from the retention base
Extrapolating a straight line past the observation window
Censored months become invented months
Report "no payback within observed window"

When does cohort analysis break?#

When entity identity is unstable. Users share logins, companies merge workspaces, contracts restart under new legal entities, marketplace participants multi-home. Establish an identity-resolution policy and preserve raw events so cohorts can be rebuilt.

When a calendar shock hits every cohort at once. A price increase, an outage, a regulatory change, or a seasonal cycle affects all ages simultaneously, and an age-aligned grid will smear it across the diagonal. Pair the cohort matrix with a calendar-time view.

When samples are small. Percentage differences between cohorts of thirty accounts are mostly noise. Report counts, intervals, and qualitative account evidence.

When contracts are long. In enterprise, logo churn arrives at renewal — possibly 24 or 36 months after the problem started. Leading indicators (adoption breadth, sponsor turnover, realised value, support severity) are more useful, but they are leading indicators, not retention, and should never be relabelled as such.

Frequently asked questions

01

Monthly or weekly cohorts?

Match the grain to the buying cycle. Self-serve products with fast activation often need weekly cohorts to see onboarding effects; enterprise contracts need quarterly. Monthly is a reasonable default, but if a cohort has fewer than about 30 members, widen the period rather than reporting noise.

02

How long until a cohort tells me something?

Long enough to clear the first churn cliff, which for most subscription products sits at the end of onboarding or the first renewal. Before that, you are measuring your onboarding funnel, not your retention. Use survival curves with censoring so partial cohorts still contribute evidence.

03

Should reactivated customers rejoin their original cohort?

Pick one policy and write it down. Returning them to the original cohort measures the relationship; treating them as a new cohort member measures the sale. Both are defensible; silently switching between them is not.

04

Do I need Kaplan–Meier, or is a simple retention rate enough?

A simple rate is fine when every account in the cohort has been observed for the full window. The moment accounts are still active at the data cutoff, or leave the panel for reasons other than churn, the naive rate is biased and the product-limit estimator is the cheap fix.

05

My NRR is 115% but revenue is flat. How?

NRR measures only the customers who existed at the start of the window. Flat total revenue with high NRR means new-customer acquisition has stalled or new cohorts are smaller than they used to be. That is precisely the diagnosis a cohort grid delivers and a blended metric hides.

Sources#

  1. E. L. Kaplan and Paul Meier, "Nonparametric Estimation from Incomplete Observations", Journal of the American Statistical Association 53(282), June 1958, 457–481. Original derivation of the product-limit (Kaplan–Meier) estimator and the correct treatment of censored observations. Open-access copy.
  2. Vertex, Inc., "Vertex Announces Fourth Quarter and Full Year 2025 Financial Results", 11 February 2026. Primary disclosure of NRR (105%), GRR (94%), ARR ($671.0m), AARPC ($137,867), and the company's written ARR/MRR and NRR definitions.
  3. Similarweb Ltd., "Similarweb Announces Fourth Quarter and Fiscal 2025 Results", 17 February 2026. Primary disclosure of overall NRR (98%) versus NRR for customers with ARR ≥ $100k (103%), customer counts (6,128 total; 454 ≥ $100k), ARR concentration (63%), and the company's ARR, NRR, CAC, and CAC-payback definitions.
  4. U.S. Securities and Exchange Commission, Commission Guidance on Management's Discussion and Analysis (Release No. 33-10751), issued 30 January 2020, effective 25 February 2020. Interpretive guidance on disclosing key performance indicators: definition, method of calculation, usefulness, management use, and methodology changes.

Source-use note: Northwind and every figure in the worked example are hypothetical. The company disclosures above are primary statements by the reporting companies about their own metric definitions; they establish that such definitions vary, not that any particular definition is correct for your business.

Author

Dr. Sarah Zou

Independent economist · EconNova

Commercial strategy for technical products, with a focus on pricing, unit economics, and the operating choices behind the model.

About Sarah

Topics

cohort analysisretentionNRRGRRsurvival analysiscensoringpaybackstartup metrics

Cite this page

Suggested citation

Zou, S. (2026). Cohort Analysis for Founders: Retention, Revenue, and Payback by Start Date. In Unit Economics. Pricing & Monetization Wiki. https://sarahzou.com/wiki/unit-economics/cohort-analysis

Open license

Reuse with attribution

This content is available for reuse. When referencing or republishing it, please credit Dr. Sarah Zou and link back to the original source.

Licensed under Creative Commons Attribution 4.0 International. You may share and adapt the material with appropriate credit.