The question nobody was measuring
The AI-visibility category has settled on a comfortable story: get recommended by ChatGPT, and the customers follow. What the category hasn’t published is the follow-up question — once an engine recommends a business, how long does that last?
Nameworthy measured 383 first-position transition pairs to produce this benchmark.
We could answer it because our measurement protocol accidentally builds the dataset: frozen questions, re-run on consecutive days, across six engines, logged every time. Pull every pair of runs where the same frozen question hit the same engine on two different days, and you can count survival directly: was yesterday’s first-named business still first today? Still present at all?
This first edition draws on 383 first-position transition pairs and over 1,800 any-presence pairs from our panel archive — multiple verticals, multiple metros, every run logged. The businesses measured were unmanaged: nobody was defending these positions. That’s the point. This is the natural decay rate — the “before” against which any managed program should be judged, including ours.
The headline number
An unmanaged #1 AI recommendation has an expected lifetime of roughly one to two days.
Per engine, the probability that the first-named business on a question is still first when the same question is re-run a day or two later:
| Engine | First-position hold rate (cross-day) | Expected lifetime at #1 |
|---|---|---|
| Perplexity | 52% (45/87) | ~2.1 days |
| ChatGPT | 50% (21/42) | ~2.0 days |
| Gemini | 37% (45/122) | ~1.6 days |
| Google AI Mode | 37% (32/86) | ~1.6 days |
| Google AI Overviews | 16% (4/25) | ~1.2 days |
| Claude | 14% (3/21)* | see note |
* Claude’s sample is the smallest and spans a mode split (local-places answers vs. web-research answers) this edition can’t fully separate from churn — treat it as a floor, not a point estimate. Expected lifetimes are geometric extrapolations from day-scale gaps: directional, not decimal-precise. We publish the raw fractions so you can check the arithmetic.
Even presence — being named anywhere in the answer, not just first — is unstable: any-presence hold rates ran from 61% (Perplexity, the stickiest surface we measured) down to the thirties and forties on Google’s AI surfaces.
Three independent checks that say the same thing
A number this stark needs corroboration, and it has it — one external, two internal:
- SISTRIX (external), analyzing 82,619 qualified prompts and 1,548,213 snapshots over 17 weeks (17 December 2025 to 8 April 2026), found ChatGPT replaces up to 74% of its cited sources every week. Week-scale source rotation is exactly what day-scale position churn compounds into.
- Our paired re-run study (internal): across paired runs of identical frozen questions, the first-named business changed 52.2% of the time.
- Our AI Overviews competitor tracking (internal): the set of competitors named on a question changed 51% cross-day — the churn isn’t just position one; the whole slate reshuffles.
Three instruments, three datasets, one conclusion: AI answers are re-drawn, not maintained.
What decays slower: the seven-day anchor
One longitudinal case in the archive is worth the whole table. A panel subject — a retailer with a published, corroborated, exactly-quotable claim on its site — held #1 on the exact-match question at day seven, while in the same seven-day window losing a generic question’s #1 down to #4. Zero competitor action in either case. Same business, same site, same week — the position anchored to a concrete published fact held; the position held by generic relevance didn’t.
That’s a single case, so we hold it lightly. But it points where the durability program goes: positions are not equally perishable, and the published, verifiable, liftable fact is the most durable substrate we’ve measured. It’s also consistent with everything else our field studies keep finding — engines hold on to what they can quote and verify.
What this means if you run a business
Monitoring is not maintenance. A dashboard that tells you weekly that your citation vanished is a subscription to bad news. The number that matters isn’t “were you recommended once” — it’s weeks held: how long a won position survives, per engine, against the one-to-two-day unmanaged baseline above. Holding requires somebody to notice the decay event and re-harden the source that slipped — that month, not next quarter.
And no one can promise you a position. This data is also why we refuse to guarantee rankings and put that refusal in our pricing: any specific position is a probabilistic outcome that churns daily. What can be honestly sold is measurement, the fix work, the re-hardening loop — and the durability record that proves whether it’s working.
Method, limitations, and what the next edition adds
Method. Every pair consists of the same frozen question, on the same engine, run on two distinct dates (median gap 1–2 days), from clean measurement vantages, drawn from our panel archive. States are scored as first-named / named / absent; a hold means the state didn’t degrade. Ongoing streaks are right-censored. Engines are reported separately, never blended.
Limitations we know about. Business-name matching is automated and imperfect, which inflates apparent churn modestly. Day-scale gaps extrapolated to lifetimes assume the decay is memoryless. Claude’s mode split needs conditioning. All of this is fixable and disclosed — the numbers here are honest fractions of logged runs, not projections.
Next edition adds true week-scale holds from monthly re-measurement cycles, hand-checked entity resolution, and — as managed data accumulates — the comparison this report exists to enable: weeks-held under management vs. the unmanaged baseline. This page will be updated in place, with the revision dated.