One name, two products
Every major AI engine ships two things that share a name and a model family but behave like different products: the consumer app a customer opens, and the developer API a tool calls. The distinction sounds technical. It decides whether a visibility number means anything.
The app is the model wrapped in a product. It browses the live web, remembers you, follows a long hidden system prompt, routes your question to whichever model version the vendor is shipping this week, and — increasingly — plugs into shopping and connectors. The API is the model with almost none of that: no web search unless the developer switches it on, no memory, no hidden prompt unless the developer writes one, and a model version the developer pins. Same engine underneath. Very different answers on top.
The measured gap
This isn’t only theoretical. In a controlled study of 900 trials published in 2026, the AI-measurement firm Petra Labs compared the brands recommended by logged-in ChatGPT against the OpenAI API on the same buying-intent prompt. The top five brands overlapped by only about a quarter — three of the app’s top five never appeared in the API’s list at all — and the differences between surfaces ran 1.5 to 3.6 times larger than the model’s own run-to-run randomness. In plain terms: the app-versus-API gap is real, and it’s bigger than the noise you’d expect from asking twice.
Petra sells measurement tools, so read the number with that in mind — but the method is disclosed, and the direction matches everything else below.
What’s different, engine by engine
The specifics vary, but the pattern holds across all four majors.
- Live web search is opt-in in the API, automatic in the app. With OpenAI, Google, and Anthropic, a developer has to switch web search on as a tool — it’s off by default — while every consumer app browses on its own. So the bare API tends to answer local and current questions from training-data memory, and the app answers them from the live web, with citations. That alone reshuffles which businesses get named.
- The app runs a hidden system prompt the API doesn’t. Anthropic makes the cleanest version of this point itself: it publishes the system prompts that run its claude.ai app and states plainly that they do not apply to the Claude API. The app you talk to and the API a tool calls are, by the vendor’s own description, different products.
- The app remembers; the API is stateless. ChatGPT, Gemini, and Claude all personalize the logged-in experience from past chats, stated preferences, and rough location. The API knows none of that unless a developer rebuilds it by hand, so it answers a generic version of every question.
- The app silently routes models; the API is pinned. ChatGPT’s product picks a fast or a “thinking” model per query and changes its default over time; Google’s AI Mode runs a custom version of Gemini that quietly shifts versions. A tool calling a pinned API model isn’t seeing the model your customer’s question was actually routed to.
- The app has a commerce layer the API can’t reach. ChatGPT now surfaces named merchants through its shopping and checkout features. Those businesses appear in the app and are invisible to a raw API call — a difference that lands squarely on whether a business gets named.
Why this quietly breaks measurement
Most automated visibility tools face a choice: measure the cheap, scriptable, stable API, or do the harder work of driving the consumer interface the way a person does. That choice determines what their numbers mean.
The good news is that the serious dedicated tools have moved to the interface — Ahrefs’ Brand Radar says it runs prompts through the public web interfaces to reflect real user experience, and Profound and Otterly both market monitoring the front-end rather than API calls. The catch is that plenty of cheaper tools and general rank-trackers still hit the API because it’s easy, and some don’t disclose which surface they use at all. When a dashboard reports your “ChatGPT visibility,” the honest question is which ChatGPT — and if the tool won’t say, the safe assumption is the one your customers never see.
What to do about it
You don’t need to become an expert in system prompts and grounding tools. You need one habit: ask any measurement — a tool, an agency, a free scan — which surface it read. The consumer app, logged out, the way a customer sees it? Or an API proxy that’s cheaper to run? If it’s the second, the number is directional at best, and worst on exactly the local, current questions that decide who gets hired.
That’s why Nameworthy measures the consumer interface on purpose, tells you which surface each result came from, and keeps every run on file so you can check the work. It’s slower than a scripted API call. It’s also the only version that matches what happens when your next customer asks. For the ChatGPT-specific version of this, see why the ChatGPT API doesn’t match the app; for the broader case on measurement tools, we pulled a rank tracker apart on the same point.