The Method

A single AI rank check is a coin flip. We don't sell coin flips.

Nameworthy measures what six AI engines say when your customers ask who to hire, then reports it honestly. Every high-intent question runs at least three times, on every surface, with the exact wording frozen and every run screenshotted. We report a rate across runs, per engine, never a single lucky answer blended into one number.

Updated July 18, 2026Reading time 6 minProtocol v1

Why isn't one check enough?

Because the same question, asked twice, gives two different answers. AI engines are non-deterministic by design: they generate a fresh set of search queries almost every time, and the businesses they name shift from run to run. Research on repeated prompting has found that only a small fraction of citations stay consistent across three runs of the identical prompt. One check tells you what happened once. It does not tell you what your customers see.

So we don't ask once. Try it — here is one real buying-intent question, run three times on the same engine, minutes apart:

Run 1. Q: best divorce lawyer in Pasadena 1. Firm A 2. Firm F 3. Firm B (sourced: Avvo, Expertise, Reddit)
Run 2. Q: best divorce lawyer in Pasadena 1. Firm F 2. Firm A — Firm B: absent — (no web search this run)
Run 3. Q: best divorce lawyer in Pasadena 1. Firm B 2. Firm A 3. Firm D (Firm B: absent → #1 in minutes)

Same engine, same question, same day. Three answers. Now imagine deciding your strategy off the first one. Anonymized from a real panel.

The retention truth behind the subscription: SISTRIX, analyzing 82,619 prompts over 17 weeks, found ChatGPT replaces up to 74% of its cited sources every week. A citation you win once is a position you have to actively hold. That is not a content project with an end date. That is maintenance, and we say so.

What does Nameworthy measure — the API or the app?

The surface your customers actually use. An API call and the consumer app can return different model versions, different search behavior, and different citations; an API request without live search returns the model's memory, not what a person sees on screen. So we collect buying-intent questions from the real consumer interface. Manual browser runs are ground truth by definition — the same standard behind the published local-search studies this field is built on.

Why measure from a clean, logged-out vantage?

Because measurement shouldn't depend on whose account is asking. Google's own AI Overviews documentation says plainly that “people's interactions with Search and these AI experiences” — what they search, what feedback they give — are used to improve its generative models, with human review in the loop. Signed-in surfaces carry history and personalization; a clean vantage answers the only question that matters: who does the engine name for a stranger? Every run in our logs is collected under that standard, and where an engine offers a memory-off mode, we use it — a discipline we've cross-checked against our own logged-in instruments to confirm the archive measures the market, not the account.

How do you score an answer?

Scoring is pre-registered — defined before the baseline runs, and never reinterpreted mid-engagement. Every run is graded against four fixed events, for your business and for each competitor:

  • Mention. Your business is named anywhere in the answer text.
  • Recommendation. You appear in a recommended list — and we record your ordinal position.
  • Citation. Your URL is linked as a source the engine drew from.
  • Sentiment. How the answer treats you: positive, neutral, or negative.

Why report every engine separately?

Because each engine runs on a different source stack, and winning one says nothing about the others. In our panels, any two engines name overlapping businesses only about 12% of the time — 1,984 runs across 139 question sets, with per-pair overlap running from 6.7% to 11.3%. Blend them into one score and you hide the exact thing you're paying to see. So we never blend. Every fix we prescribe is aimed at a named engine.

Why these six engines, and not ten?

Some tools advertise coverage of ten or more engines. We measure six on purpose: the surfaces where local, commercial-intent questions actually get asked and answered — ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews, and Google AI Mode — each measured in the consumer app, three runs per question. A wider engine list measured shallowly, through APIs that don't match what customers see, is a logo wall, not a measurement. When another surface earns real commercial-intent share, we add it the same way we measure the rest: app, repeated runs, per-engine report. Until then we'd rather be right about six than decorative about twelve.

What's the statistics behind it?

A citation rate is a proportion, so we treat it like one. We report the rate across runs with a confidence interval, not a bare percentage. Repeated-generation measurement with panel-level means is now the published standard for evaluating language models; the variance research is clear that single-run estimates swing by several points even at temperature zero, and that prompt wording is a dominant source of error. That last finding is why the panel is frozen: we never edit a prompt mid-engagement, because changing the words changes the number.

Read next

The research — how AI engines actually source local-business recommendations.

Ready to see yours

Get your free report — your customers' real questions, the major engines, 48 hours.

Common questions

How many AI engines do you measure?
Six, reported separately: ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews, and Google AI Mode. They run on different sources and disagree too much to average.
How many times do you run each question?
More than once, always, because a single run is a coin flip. AI answers are non-deterministic — the recommended businesses change from run to run — so we report a rate across runs, not one observation. The 10 highest-intent questions in a panel run three times each on Gemini, Perplexity, AI Overviews and AI Mode and twice on the ChatGPT and Claude apps, where every session is driven by hand rather than scripted. The 15 supporting questions run once on the four automated engines. That is 220 scored observations per panel, and every one of them is in the log we hand you.
Do you measure the API or the app?
The consumer app — the surface your customers actually use. API responses can differ in model version, search behavior, and citations, so we collect buying-intent questions from the real interface and document how.
Can you guarantee I'll show up in ChatGPT?
No, and anyone who guarantees a specific placement is not being honest — positions churn week to week. What we guarantee is presence: named by at least one AI engine on at least 5 of your 25 client-approved questions by day 90, verified in the run log we hand you, or month four is free.

The report is free

See your numbers, measured the honest way.

Twenty-five questions, six engines, forty-eight hours, every run on file.

Get your AI visibility report