How to Read a SaaS Benchmark

One of the things that excites me in a way that demonstrates just how uncool and unhip I am like no other is business performance metrics. And no business has more jargon-laden, misunderstood, and vague metrics than SaaS. Which makes it the most fun. However, with SaaS still being the hot thing that everyone wants to disrupt, innovate, or simply cash in on, its metrics have become something of a vision board for whatever the person citing said metrics wants them to present.
For example, two articles about B2B SaaS churn, published four months apart in 2026, both on the first page of Google. The first says the median churn is 3.5% per month. The second also says median churn is 3.5% per year. Same number. Same metric. Same industry. A twelve-fold difference in what it actually means.
One describes a business shedding roughly a third of its revenue base every year. The other describes a business losing almost none of it. Anyone setting a retention target off either page is setting it off a figure that someone typed into a different box than the one it came out of. Yet every year you’ll see people cite these figures as objective benchmarks that hold true across the entire industry when in reality, they barely mean what we claim they do.
Out of curiosity I went looking for the real number. There isn’t one. At least not one you can reliably tie back to any kind of credible data or research. It’s all just analyst firms guesstimating where they think it should fall or anecdotal callouts. The reason why actually tells us a lot about the state of the industry we call “SaaS”.
The number you cannot look up
Take net revenue retention, the metric most SaaS boards treat as the best single proxy for product health. Here is what the credible sources say right now.
SaaS Capital, surveying more than 1,000 private B2B SaaS companies in the $3M–$20M ARR band, puts median NRR at 103%, essentially flat year over year. Jamin Ball’s Clouded Judgement, tracking the public software index, puts it at 108%. A separate Ball analysis cited by SaaStr describes public NRR as “still over 110%.” Benchmarkit, whose panel includes AI-native companies alongside traditional SaaS, reports gross revenue retention falling from 88% to 84%. ChartMogul, categorizing roughly 3,500 software companies, puts median NRR for AI-native companies specifically at 48%.
Retention, 2026
Five credible sources. Five answers. None of them wrong.
There are no bars in this figure, deliberately. Four of these measure net revenue retention and one measures gross, and they are drawn from populations that barely overlap. Putting them on a shared scale would commit the exact error this article is about. The line that decides what each number means is the one beside it, not the number itself.
Every one of those is real research from an organization that publishes its methodology. None is wrong. They do not reconcile.
CAC payback is worse, because the spread is wider and the folklore is stronger. Most operators still carry a rule of thumb somewhere between twelve and eighteen months. Data-Mania puts the 2026 median at 15 months. Aleph and Benchmarkit, from a 342-company survey, put it at 16. G-Squared’s CFO commentary puts it at 18, ranging 8 to 24 by deal size.
Clouded Judgement puts the median for public software companies at 36 months.
One source says payback takes a year and a quarter. Another says three years. Same metric, same year, both carefully measured.
CAC payback, 2026 medians
The same metric, measured on two different populations
Bars are proportional to months, on a shared scale, because unlike the figure above these do measure the same thing. Every bar is labeled with its population, so the two colours are never the only thing telling them apart.
This is where most benchmark posts stop and pretend the discrepancy is a rounding error you can average away. It isn’t.
Why they disagree
Every one of those gaps has a cause, and the cause is the information underneath.
The 36-versus-16 gap is public versus private. Public software companies are larger, sell to enterprises, carry heavier sales organizations, and spent a decade inside growth-at-any-cost incentives that private bootstrapped companies never had access to. A 342-company survey of mostly private SaaS is not measuring the same population as an index of publicly traded software. They are counting different companies, and both numbers are right for the companies they counted.
The 108-versus-48 gap is traditional versus AI-native. ChartMogul’s AI-native cohort retains at less than half the rate of traditional B2B SaaS, and their data shows exactly where it breaks. AI products priced above $250 a month hold 70% gross retention, roughly at parity with normal SaaS. Between $50 and $249, that falls to 45%. Below $50, it collapses to 23%. They call it the tourist problem: cheap AI tools get tried and abandoned, expensive ones get deployed and kept.
AI-native gross retention, by price point
Cheap AI tools get tried and abandoned. Expensive ones get deployed and kept.
One hue, darker where retention is higher, because the price bands are an ordered scale rather than separate categories. Source: ChartMogul, primary period running through September 2025.
The 91-versus-84 gross retention gap — and this is the one people miss — is sample composition. SaaS Capital’s panel is bootstrapped, small, traditional. Benchmarkit’s explicitly includes AI-native companies. Fold a cohort retaining at 40% GRR into a panel and the median moves without a single company in either group changing behavior.
So “what is the SaaS retention benchmark” has stopped parsing as a question. “SaaS company” is no longer one population. A bootstrapped $8M-ARR vertical tool, a public enterprise platform, and an AI-native product selling $250-a-month seats have different physics. Averaging them produces a number that describes none of them.
That is a category failure rather than a measurement failure, and it is the most consequential thing happening in SaaS benchmarking right now.
The layer underneath
There is a second problem, and it is the one that will actually cost you money.
Search any SaaS metric benchmark and the first page fills with a near-identical title pattern: “[Metric] Benchmarks 2026: [oddly precise number] ([N] Companies).” The publishers are companies you have not heard of. I pulled several and read their sourcing rather than their conclusions.
The churn page claiming 500-plus companies has no methodology section. It does not say where the 500 companies came from, what period they cover, how they were selected, or what was measured. It carries numbered footnotes running to thirteen references and no bibliography. The markers appear in the text with nothing at the bottom to resolve them against. The named sources that do surface in the prose are software tools and vendors, one of them credited to the co-founder of another site in the same genre. It is bylined to the company itself rather than a person, and it was published in April 2026 and modified in July, which is the only verifiable fact on the page.
A second page, from an unrelated publisher, carries a “Research Team” byline, states no sample size and no methodology, and cites as its sources two other pages of the same type.
That is the mechanism. These pages cite each other. A figure gets published without provenance, the next page picks it up as a citation, and by the third hop it has three sources and no origin. Nothing in the chain ever touched a company’s books. The 3.5% appearing as a monthly rate on one page and an annual rate on another is what that looks like when the copying goes slightly wrong. Nobody caught it because there was no underlying dataset to check it against.
I want to be precise about the accusation, because it does not apply to everyone writing benchmark content. G-Squared’s piece is authored by a named CFO, cites Benchmarkit, High Alpha, Bessemer and SaaS Capital by name, and makes no claim to original data. That is honest synthesis, and it is useful. SaaS Capital states its sample, its ARR band and its survey cadence. ChartMogul publishes its categorization method and its cutoff. Clouded Judgement is one analyst’s index, openly labeled as such.
The distinction is not whether a page ran its own study. It is whether the page tells you where its numbers came from. Honest synthesis does. The other kind manufactures the appearance of rigor, with a specific sample size in the headline, footnote markers in the body and a confident median in bold, and none of the substance behind it. That is harder to spot precisely because it looks more like research than an honest summary does.
Why this got worse in 2026
Benchmark content has always been soft. Two things happened this year that turned soft into unusable.
The first is the category fracture described above. Until recently a “B2B SaaS company” was a reasonably coherent thing: per-seat pricing, annual contracts, 70 to 80 percent gross margins, land and expand. A median across that population meant something because the population was similar. That coherence is gone. AI-native companies now sell at price points and retention profiles that look nothing like traditional SaaS, and they are being folded into the same panels. The credible publishers diverged because the thing they were all measuring stopped being one thing.

The second is supply. Producing a plausible benchmark article used to require either original research or the patience to assemble other people’s. It now requires a prompt. The result is a layer of content that is cheap to produce, optimized to rank, and structurally incapable of being right, because there is no dataset underneath it to be right about. It fills the first page of results for exactly the queries an operator types when they want a number fast.
Those two forces compound. Real sources disagree more than they used to, which creates genuine uncertainty. Fabricated sources resolve that uncertainty with false confidence and better SEO. The operator searching for “SaaS churn benchmark 2026” gets the confident wrong answer before the careful qualified one.
What this costs you
The obvious cost is a wrong target. If your board sets a twelve-month CAC payback goal because that is the number on the slide, and your real peer set runs at 24 to 36, you will spend a year failing an exam nobody else is sitting.
The expensive cost is allocation. Benchmarks do not just measure, they direct money. A retention figure that looks bad against a fabricated median gets a remediation program, a headcount and a quarter of executive attention. A CAC payback that looks acceptable against a soft benchmark gets left alone when it was the thing you should have fixed. Both errors are invisible, because the benchmark that caused them is never revisited once the decision is made.
I have sat in the meetings where one slide of “industry standard” figures redirected real budget, and nobody in the room could have said where the figures came from, because the deck did not say and nobody asked. The risk is not that a number is wrong. It is that a wrong number is load-bearing.
Five questions before you use one
Benchmarks are not useless. They need reading like any other evidence. Five questions, ordered by how fast they disqualify a source.
1. Who was counted, and how many? A benchmark with no stated sample is an assertion wearing a lab coat. “500+ companies” in a headline with no methodology in the body is worse than no number at all, because it borrows the authority of research without doing any.
2. Public or private, and at what size? This one question resolves most of the CAC payback confusion. Thirty-six months and sixteen months are both correct for their populations. Establish which population you are in before you pick a target. If the source does not say, you cannot know.
3. Does the panel include AI-native companies? In 2026 this moves retention figures more than any other variable. If the report does not disclose it, its retention numbers are unusable, because you cannot tell whether you are looking at a cohort that includes companies churning at 60% a year.
4. Where does one specific figure trace to? Not “does it have citations.” Follow one. Take the number that matters most to your decision and chase it upstream. If two hops land you on another benchmark page rather than a survey, a filing, an earnings call or a named index, stop using the source entirely. This takes about ninety seconds and it disqualifies more sources than the other four questions combined. When I traced the 3.5% churn figure, the first hop led to a page citing three tools by name and no study; the second led to a page citing the first. There was no third hop, because there was nothing underneath.
5. What is the date on the underlying data, not the page? Pages get republished with the current year in the title. ChartMogul’s AI-native retention data is the best available on that cohort and its primary period runs through September 2025. Useful, a year old, and worth citing as exactly that. A “2026 benchmarks” page built on 2024 data is not fresh because someone edited the headline.
The version that works
Use benchmarks as range and direction, not as targets.
The useful reading of the retention data above is not a number at all. It is this: private small-cap SaaS retention is roughly flat, public retention has declined every quarter since late 2022 and now sits below its pre-pandemic level, and AI-native retention is dramatically worse and sharply stratified by price point. Every one of those statements is supportable from named sources with published methods. None requires pretending the five figures agree.
Then build your own comparison set. Ten to fifteen companies you can name, matched on model, ACV and buyer, beats an industry median every time, because you can explain the differences. Public comparables file their numbers quarterly. Private ones surface in investor updates, earnings calls and the occasional honest founder post. That set will be small and specific and unglamorous, and it will tell you more than any median, because you will know exactly who is in it and why each one is there.
Keep the industry figures for the questions they can actually answer: is this metric improving or deteriorating across the market, and roughly how far from the middle am I. Those are real questions and benchmarks are good at them. “Am I hitting the standard” is not a question the current data can answer, because there is no standard. There are five, and they disagree by more than most companies’ annual improvement.
And when someone puts an industry benchmark on a slide in your company, ask where it came from. Not to be difficult, but because if the answer takes longer than a minute to produce, you have learned something important about the decision you are being asked to make.
Sources: SaaS Capital private B2B SaaS survey (>1,000 companies, $3M–$20M ARR, median NRR 103%); Clouded Judgement public software index (NRR 108%, “still over 110%” via SaaStr, CAC payback 36 months); Benchmarkit GRR 88% to 84% panel including AI-native; ChartMogul ~3,500 companies categorization, AI-native median NRR 48%, GRR stratified >$250 70%, $50-$249 45%, <$50 23% “tourist problem”; Data-Mania 2026 median CAC payback 15 months; Aleph + Benchmarkit 342-company survey median 16 months; G-Squared CFO commentary 18 months range 8-24 by deal size; 3.5% monthly vs annual churn example from 2026 SERP pages (500+ companies no methodology, 13 footnotes no bibliography, Research Team byline).
Featured image generated with AI.
Written by
Ryan Frazier
He’s spent 18 years building and leading marketing teams, from Series A startups to multi-billion-dollar public companies — four of them scaled past the $50M, $100M and $250M ARR marks, and all four through to acquisition. He writes The Positioning, on why winning has less to do with being right than with being well-positioned at the convergence of time, place, and resource.
More about Ryan →