What we did
There’s a widely repeated claim that AI-visibility tools querying the OpenAI API don’t measure what a customer sees in the ChatGPT app. We wanted our own numbers, on the questions that actually matter — local, buying-intent questions like “best [service] near me.”
So we ran a controlled comparison. We took 40 real buying-intent questions across four local service verticals — personal-injury law, family law, med spa, and cannabis retail, one market each in four U.S. states — and asked every question three ways:
- The ChatGPT app — the logged-in consumer product, with web search on (the surface a customer actually uses).
- The bare API — the OpenAI API with no browsing tool, on the same underlying model the app runs (this is how the cheap, scriptable tools query “ChatGPT”).
- The API with web search — the same model, browsing enabled, to isolate how much of the gap is just the ability to search.
We ran each question three times per surface, recorded the businesses each one recommended, and measured the overlap. (Four of the 40 were branded reputation questions about a specific business; those are set aside from the head-to-head, leaving 36 comparison questions.) Because the app and the browsing API arm run the same model, the differences below come from the product wrapper around it — a hidden system prompt, live browsing, memory — not from a different brain.
The headline
Across all four verticals, here’s how often each API surface recommended the same businesses as the app:
| API surface | Overlap with the app | Recovered the app’s picks | Named any local business |
|---|---|---|---|
| Bare API (no browsing) | ~8% | 17% | 72% |
| API with web search | ~16% | 36% | 97% |
| Search-native model | ~9% | 32% | 100% |
For reference, the app named a local business on 97% of these questions. “Overlap” here is the share of all the businesses recommended across both surfaces that both surfaces recommended — so 8% means that, of everything the app and the bare API named, only about one in twelve was named by both.
What the bare model recommends instead
When the API can’t search, it recommends from memory. And memory favors the businesses with the loudest national footprint — the biggest advertisers and the national chains — over the local specialists the app surfaces after browsing. In two of our four verticals, the bare API’s recommendations overlapped the app’s by essentially zero: it named a completely different, nationally-familiar set, while the app named local, credential-verified businesses it found by searching in the moment.
Browsing narrows the gap. It doesn’t close it.
Give the same model a web-search tool and it moves toward the app: overlap roughly doubles, and it recovers about a third of the app’s picks. That’s the expected result — browsing is the single biggest thing the bare API is missing. But “a third of the picks” is still a different answer. The app’s wrapper does more than search; it also applies a long hidden system prompt and personalization that the API doesn’t, and those keep the two surfaces apart even when both are browsing.
A “search-native” API isn’t the shortcut either
We also tested a dedicated always-searching model — the kind of search-first API some tools lean on. It named a business every single time, but its overlap with the app was only about 9% — barely above the bare model. The reason is instructive: it browses widely and returns long, noisy lists, so it names more, but not the same. Always searching is not the same as matching what a careful consumer product recommends.
It even names businesses the app never surfaced
The gap runs both ways. On several questions, an API surface recommended a business that the app never named for the same question — including, in a few cases, the very business we were separately measuring, which the app had left out of its shortlist. That’s the point worth sitting with: the API isn’t a lower-resolution version of the app. It’s a different answer, and sometimes it’s confidently pointing at businesses the app quietly passed over.
The vertical spread
The gap wasn’t uniform. It was smallest in the vertical where the market’s most prominent name is both nationally famous and locally dominant — there, the bare model’s memory happened to line up with reality more often. It was near-zero in the verticals built on hyper-local specialists, where the bare model reverts to national chains it can recall. That’s a useful warning: how badly the API misrepresents your market depends on your market, and you can’t know which case you’re in without measuring.
What this means if a tool measures your “ChatGPT visibility”
Most automated tools query the API because it’s cheap and stable. Our data says that number describes a surface your customers don’t use — and on local, buying-intent questions, it diverges most exactly where it matters. If a dashboard tells you where you stand in “ChatGPT,” the honest first question is which ChatGPT — the browsing, personalized app a customer opens, or an API proxy that’s easier to script. That’s why we measure the consumer app, in a fresh session, and tell you which surface each result came from.
How to read this, honestly
This is a first cut, and we’ll say so plainly. It’s four verticals, one market and one focal business each, 36 comparison questions, three runs per surface — enough to establish the shape of the gap, not to put a decimal point on it. We measured the overlap of full recommendation sets, which is a stricter test than comparing only the top few names. And it’s our own study, run to inform our own method — read it as directional evidence that the API and the app are different products, which is the same conclusion the broader, cross-engine case reaches from the vendors’ own documentation.
If you want to see it for your own business — measured on the app a customer actually uses, across all six engines — that’s the free report.