Person
Person

Sep 14, 2026

By Victor Teran

Why AI Products Churn Harder Than Software Did

AI products are the easiest thing in the world to try once and the easiest thing in the world to stop using. That combination produces a shape of churn traditional software never had, and the metrics most founders watch are blind to it.

Retention

AI

Churn

An AI product gets a benefit no SaaS tool ever got: people will try it out of curiosity. Signup is cheap, the demo is impressive, and the first session usually delivers something that feels like magic.

Then a large share of those people never come back, and the dashboard does not scream, because acquisition looks excellent and the first-session numbers look excellent too. The problem sits between them, in the gap where a habit was supposed to form and did not.

a16z named the measurement problem directly in September 2025. Rather than a single retention number, they argue the useful benchmark for AI companies is the ratio of month-twelve to month-three retention, which isolates how customers who survive the initial curiosity churn behave over their first full year (Rodriguez and Immerman, a16z). The phrase they use for the people who leave early is tourist churn, and it is the right mental model.

Key takeaways

  • AI products convert curiosity into signups extremely well, which inflates every top-of-funnel metric and hides the retention problem for months.

  • a16z's proposed benchmark for AI companies is the month-twelve to month-three retention ratio, precisely because month-one numbers are distorted by curiosity traffic.

  • The failure is usually not the model. It is that the product produced an impressive output once and never became part of a repeated workflow.

  • Cohort your curiosity users separately from your intent users, or the two populations will average into a number that describes nobody.

Why do AI products get such a flattering first month?

Because trying an AI product costs almost nothing and often produces something worth screenshotting.

Traditional software asked for setup. Import your data, invite your team, configure the thing, and only then find out whether it helps. That friction filtered hard: the people who finished setup were people with a real problem.

AI products removed the filter. Type a prompt, get a result. That is wonderful for signup volume and it means your funnel now contains two populations that behave nothing alike: people with a job to do, and people who wanted to see what it does.

Averaged together they produce a month-one retention figure that describes neither. The curiosity group leaves regardless of product quality. The intent group's behaviour, the only signal that predicts a business, is buried underneath them.

This is why a16z's month-twelve over month-three framing is more than a metrics preference. It is an attempt to look at the population that was ever going to matter.

What actually breaks between the demo and the habit?

The product proved it could do something impressive, and never became the thing you reach for on a Tuesday.

A demo is a single output. A habit is a trigger, an action, and a reward that repeats. Most AI products nail the middle and neglect the other two.

The failure modes we see repeatedly in audits:

What is missing

What it looks like in the data

No trigger in the user's existing workflow

Strong first session, then nothing. Users never return unprompted

Output impressive but not trusted

Users generate, then abandon before using the result anywhere

Value moment is a one-off, not a loop

Excellent day one, cliff by day three, flat zero after week one

No second use case

Users repeat one narrow task, never expand, churn when it ends

None of these are model quality problems, and none of them are fixed by a better model. They are product and behaviour problems, which is a relief, because they are the kind you can actually change in weeks.

FREE GUIDE

The 4 product leaks costing you growth

A short audit guide for founders. Find the four places your product leaks revenue, and what to fix first.

How do you separate curiosity from intent?

Cohort on a behaviour that a curious visitor would never bother to perform.

The specific behaviour differs by product, but the test is always the same: what would someone do only if they intended to use this for real? Connecting a data source. Inviting a colleague. Saving or exporting an output into somewhere it does real work. Returning on a second calendar day without a prompt from you.

For reference, the 2026 onboarding benchmark defines activation as reaching a defined value moment inside a seven or fourteen day window, with median B2B SaaS at 38% (Perspective AI). That window is the right discipline here too. Split your cohorts on that action and look at the two curves separately. In practice the intent curve is far healthier than the blended number suggested, and the curiosity curve is far worse. Both facts are useful and the average was hiding them.

This is the same discipline as defining an activation event, applied to a product where the population is unusually mixed. If you have not defined one, that comes first, and we have written how to do it in one sentence.

Which metric should a founder actually watch?

Return rate among users who reached the value moment, measured on a fixed window, cohorted by week.

Not signups. Not sessions. Not total generations, which is the vanity metric of this product category, because a single enthusiastic user can carry it while the base rots underneath. Activation is the moment a user first experiences core value (Amplitude), and for an AI product that is rarely the first impressive output.

Three things worth putting on one screen:

1. Activation rate among new signups, on a defined value moment inside seven or fourteen days. 2. Week-four return rate for activated users only, by weekly cohort, so you can see whether product changes moved anything. 3. The ratio of a late-month retention figure to an early one, which is the a16z point: the shape matters more than the level when your early numbers are polluted.

If the curve flattens, you have something. If it decays toward zero no matter how good month one looks, more acquisition will make the problem larger and more expensive, not smaller.

What do you do once you can see it?

Find the earliest place the intent cohort stops, and fix the workflow rather than the model.

The pattern is consistent enough to plan around. Users who return have almost always done something in their first session that connected the output to work they already had: put it somewhere, shared it, built on it. Users who do not return admired the output and closed the tab.

So the work is rarely "make the model better". It is making the first result land somewhere the user was already going to be. That is a product decision, an onboarding decision, and sometimes an integration decision, and it is exactly the kind of thing a funnel plus session replays will point at within days.

We have written separately about why churn is almost never random, and AI products are the strongest version of that argument. The users leaving look random only because two different populations are being counted as one.

WHAT NEXT

Want this fixed in your product, not just explained?

Latest Updates

(OTU® — 01)

©2026

Only 4.6% of Apps Reach $10K MRR. What the Other 95% Get Wrong

Retention

80% of App Users Churn in Week One. Here Is Where They Actually Leave

Retention

What Is an Activation Event? Define Yours in One Sentence

Activation

Person
Person

Sep 14, 2026

By Victor Teran

Why AI Products Churn Harder Than Software Did

AI products are the easiest thing in the world to try once and the easiest thing in the world to stop using. That combination produces a shape of churn traditional software never had, and the metrics most founders watch are blind to it.

Retention

AI

Churn

An AI product gets a benefit no SaaS tool ever got: people will try it out of curiosity. Signup is cheap, the demo is impressive, and the first session usually delivers something that feels like magic.

Then a large share of those people never come back, and the dashboard does not scream, because acquisition looks excellent and the first-session numbers look excellent too. The problem sits between them, in the gap where a habit was supposed to form and did not.

a16z named the measurement problem directly in September 2025. Rather than a single retention number, they argue the useful benchmark for AI companies is the ratio of month-twelve to month-three retention, which isolates how customers who survive the initial curiosity churn behave over their first full year (Rodriguez and Immerman, a16z). The phrase they use for the people who leave early is tourist churn, and it is the right mental model.

Key takeaways

  • AI products convert curiosity into signups extremely well, which inflates every top-of-funnel metric and hides the retention problem for months.

  • a16z's proposed benchmark for AI companies is the month-twelve to month-three retention ratio, precisely because month-one numbers are distorted by curiosity traffic.

  • The failure is usually not the model. It is that the product produced an impressive output once and never became part of a repeated workflow.

  • Cohort your curiosity users separately from your intent users, or the two populations will average into a number that describes nobody.

Why do AI products get such a flattering first month?

Because trying an AI product costs almost nothing and often produces something worth screenshotting.

Traditional software asked for setup. Import your data, invite your team, configure the thing, and only then find out whether it helps. That friction filtered hard: the people who finished setup were people with a real problem.

AI products removed the filter. Type a prompt, get a result. That is wonderful for signup volume and it means your funnel now contains two populations that behave nothing alike: people with a job to do, and people who wanted to see what it does.

Averaged together they produce a month-one retention figure that describes neither. The curiosity group leaves regardless of product quality. The intent group's behaviour, the only signal that predicts a business, is buried underneath them.

This is why a16z's month-twelve over month-three framing is more than a metrics preference. It is an attempt to look at the population that was ever going to matter.

What actually breaks between the demo and the habit?

The product proved it could do something impressive, and never became the thing you reach for on a Tuesday.

A demo is a single output. A habit is a trigger, an action, and a reward that repeats. Most AI products nail the middle and neglect the other two.

The failure modes we see repeatedly in audits:

What is missing

What it looks like in the data

No trigger in the user's existing workflow

Strong first session, then nothing. Users never return unprompted

Output impressive but not trusted

Users generate, then abandon before using the result anywhere

Value moment is a one-off, not a loop

Excellent day one, cliff by day three, flat zero after week one

No second use case

Users repeat one narrow task, never expand, churn when it ends

None of these are model quality problems, and none of them are fixed by a better model. They are product and behaviour problems, which is a relief, because they are the kind you can actually change in weeks.

FREE GUIDE

The 4 product leaks costing you growth

A short audit guide for founders. Find the four places your product leaks revenue, and what to fix first.

How do you separate curiosity from intent?

Cohort on a behaviour that a curious visitor would never bother to perform.

The specific behaviour differs by product, but the test is always the same: what would someone do only if they intended to use this for real? Connecting a data source. Inviting a colleague. Saving or exporting an output into somewhere it does real work. Returning on a second calendar day without a prompt from you.

For reference, the 2026 onboarding benchmark defines activation as reaching a defined value moment inside a seven or fourteen day window, with median B2B SaaS at 38% (Perspective AI). That window is the right discipline here too. Split your cohorts on that action and look at the two curves separately. In practice the intent curve is far healthier than the blended number suggested, and the curiosity curve is far worse. Both facts are useful and the average was hiding them.

This is the same discipline as defining an activation event, applied to a product where the population is unusually mixed. If you have not defined one, that comes first, and we have written how to do it in one sentence.

Which metric should a founder actually watch?

Return rate among users who reached the value moment, measured on a fixed window, cohorted by week.

Not signups. Not sessions. Not total generations, which is the vanity metric of this product category, because a single enthusiastic user can carry it while the base rots underneath. Activation is the moment a user first experiences core value (Amplitude), and for an AI product that is rarely the first impressive output.

Three things worth putting on one screen:

1. Activation rate among new signups, on a defined value moment inside seven or fourteen days. 2. Week-four return rate for activated users only, by weekly cohort, so you can see whether product changes moved anything. 3. The ratio of a late-month retention figure to an early one, which is the a16z point: the shape matters more than the level when your early numbers are polluted.

If the curve flattens, you have something. If it decays toward zero no matter how good month one looks, more acquisition will make the problem larger and more expensive, not smaller.

What do you do once you can see it?

Find the earliest place the intent cohort stops, and fix the workflow rather than the model.

The pattern is consistent enough to plan around. Users who return have almost always done something in their first session that connected the output to work they already had: put it somewhere, shared it, built on it. Users who do not return admired the output and closed the tab.

So the work is rarely "make the model better". It is making the first result land somewhere the user was already going to be. That is a product decision, an onboarding decision, and sometimes an integration decision, and it is exactly the kind of thing a funnel plus session replays will point at within days.

We have written separately about why churn is almost never random, and AI products are the strongest version of that argument. The users leaving look random only because two different populations are being counted as one.

WHAT NEXT

Want this fixed in your product, not just explained?

Latest Updates

(OTU® — 01)

©2026

Only 4.6% of Apps Reach $10K MRR. What the Other 95% Get Wrong

Retention

80% of App Users Churn in Week One. Here Is Where They Actually Leave

Retention

What Is an Activation Event? Define Yours in One Sentence

Activation

Person
Person

Sep 14, 2026

By Victor Teran

Why AI Products Churn Harder Than Software Did

AI products are the easiest thing in the world to try once and the easiest thing in the world to stop using. That combination produces a shape of churn traditional software never had, and the metrics most founders watch are blind to it.

Retention

AI

Churn

An AI product gets a benefit no SaaS tool ever got: people will try it out of curiosity. Signup is cheap, the demo is impressive, and the first session usually delivers something that feels like magic.

Then a large share of those people never come back, and the dashboard does not scream, because acquisition looks excellent and the first-session numbers look excellent too. The problem sits between them, in the gap where a habit was supposed to form and did not.

a16z named the measurement problem directly in September 2025. Rather than a single retention number, they argue the useful benchmark for AI companies is the ratio of month-twelve to month-three retention, which isolates how customers who survive the initial curiosity churn behave over their first full year (Rodriguez and Immerman, a16z). The phrase they use for the people who leave early is tourist churn, and it is the right mental model.

Key takeaways

  • AI products convert curiosity into signups extremely well, which inflates every top-of-funnel metric and hides the retention problem for months.

  • a16z's proposed benchmark for AI companies is the month-twelve to month-three retention ratio, precisely because month-one numbers are distorted by curiosity traffic.

  • The failure is usually not the model. It is that the product produced an impressive output once and never became part of a repeated workflow.

  • Cohort your curiosity users separately from your intent users, or the two populations will average into a number that describes nobody.

Why do AI products get such a flattering first month?

Because trying an AI product costs almost nothing and often produces something worth screenshotting.

Traditional software asked for setup. Import your data, invite your team, configure the thing, and only then find out whether it helps. That friction filtered hard: the people who finished setup were people with a real problem.

AI products removed the filter. Type a prompt, get a result. That is wonderful for signup volume and it means your funnel now contains two populations that behave nothing alike: people with a job to do, and people who wanted to see what it does.

Averaged together they produce a month-one retention figure that describes neither. The curiosity group leaves regardless of product quality. The intent group's behaviour, the only signal that predicts a business, is buried underneath them.

This is why a16z's month-twelve over month-three framing is more than a metrics preference. It is an attempt to look at the population that was ever going to matter.

What actually breaks between the demo and the habit?

The product proved it could do something impressive, and never became the thing you reach for on a Tuesday.

A demo is a single output. A habit is a trigger, an action, and a reward that repeats. Most AI products nail the middle and neglect the other two.

The failure modes we see repeatedly in audits:

What is missing

What it looks like in the data

No trigger in the user's existing workflow

Strong first session, then nothing. Users never return unprompted

Output impressive but not trusted

Users generate, then abandon before using the result anywhere

Value moment is a one-off, not a loop

Excellent day one, cliff by day three, flat zero after week one

No second use case

Users repeat one narrow task, never expand, churn when it ends

None of these are model quality problems, and none of them are fixed by a better model. They are product and behaviour problems, which is a relief, because they are the kind you can actually change in weeks.

FREE GUIDE

The 4 product leaks costing you growth

A short audit guide for founders. Find the four places your product leaks revenue, and what to fix first.

How do you separate curiosity from intent?

Cohort on a behaviour that a curious visitor would never bother to perform.

The specific behaviour differs by product, but the test is always the same: what would someone do only if they intended to use this for real? Connecting a data source. Inviting a colleague. Saving or exporting an output into somewhere it does real work. Returning on a second calendar day without a prompt from you.

For reference, the 2026 onboarding benchmark defines activation as reaching a defined value moment inside a seven or fourteen day window, with median B2B SaaS at 38% (Perspective AI). That window is the right discipline here too. Split your cohorts on that action and look at the two curves separately. In practice the intent curve is far healthier than the blended number suggested, and the curiosity curve is far worse. Both facts are useful and the average was hiding them.

This is the same discipline as defining an activation event, applied to a product where the population is unusually mixed. If you have not defined one, that comes first, and we have written how to do it in one sentence.

Which metric should a founder actually watch?

Return rate among users who reached the value moment, measured on a fixed window, cohorted by week.

Not signups. Not sessions. Not total generations, which is the vanity metric of this product category, because a single enthusiastic user can carry it while the base rots underneath. Activation is the moment a user first experiences core value (Amplitude), and for an AI product that is rarely the first impressive output.

Three things worth putting on one screen:

1. Activation rate among new signups, on a defined value moment inside seven or fourteen days. 2. Week-four return rate for activated users only, by weekly cohort, so you can see whether product changes moved anything. 3. The ratio of a late-month retention figure to an early one, which is the a16z point: the shape matters more than the level when your early numbers are polluted.

If the curve flattens, you have something. If it decays toward zero no matter how good month one looks, more acquisition will make the problem larger and more expensive, not smaller.

What do you do once you can see it?

Find the earliest place the intent cohort stops, and fix the workflow rather than the model.

The pattern is consistent enough to plan around. Users who return have almost always done something in their first session that connected the output to work they already had: put it somewhere, shared it, built on it. Users who do not return admired the output and closed the tab.

So the work is rarely "make the model better". It is making the first result land somewhere the user was already going to be. That is a product decision, an onboarding decision, and sometimes an integration decision, and it is exactly the kind of thing a funnel plus session replays will point at within days.

We have written separately about why churn is almost never random, and AI products are the strongest version of that argument. The users leaving look random only because two different populations are being counted as one.

WHAT NEXT

Want this fixed in your product, not just explained?

Latest Updates

©2026

Only 4.6% of Apps Reach $10K MRR. What the Other 95% Get Wrong

Retention

80% of App Users Churn in Week One. Here Is Where They Actually Leave

Retention

What Is an Activation Event? Define Yours in One Sentence

Activation