Briefing

The API is not the app: what that means for measuring AI visibility

For every major engine, the developer API and the consumer app are different products — the app browses the live web, remembers you, and runs a hidden system prompt the API doesn't. A 2026 study found the top brands ChatGPT recommends in the app versus its API overlap by only about a quarter. So measuring the API and calling it "what AI says" describes a surface your customer never opens.

Updated July 27, 2026Reading time 6 min

One name, two products

Every major AI engine ships two things that share a name and a model family but behave like different products: the consumer app a customer opens, and the developer API a tool calls. The distinction sounds technical. It decides whether a visibility number means anything.

The app is the model wrapped in a product. It browses the live web, remembers you, follows a long hidden system prompt, routes your question to whichever model version the vendor is shipping this week, and — increasingly — plugs into shopping and connectors. The API is the model with almost none of that: no web search unless the developer switches it on, no memory, no hidden prompt unless the developer writes one, and a model version the developer pins. Same engine underneath. Very different answers on top.

The measured gap

This isn’t only theoretical. In a controlled study of 900 trials published in 2026, the AI-measurement firm Petra Labs compared the brands recommended by logged-in ChatGPT against the OpenAI API on the same buying-intent prompt. The top five brands overlapped by only about a quarter — three of the app’s top five never appeared in the API’s list at all — and the differences between surfaces ran 1.5 to 3.6 times larger than the model’s own run-to-run randomness. In plain terms: the app-versus-API gap is real, and it’s bigger than the noise you’d expect from asking twice.

Petra sells measurement tools, so read the number with that in mind — but the method is disclosed, and the direction matches everything else below.

What’s different, engine by engine

The specifics vary, but the pattern holds across all four majors.

The exception that proves the rule — Perplexity. Perplexity’s Sonar API was built search-first: it browses and returns citations by default, so it behaves far more like its app than the others do. The residual gaps are memory, cross-vendor model routing, and its research workspaces — not the core question of whether it searches the web. When a claim about “the API” gets sweeping, Perplexity is the counterexample to keep it honest. The gap is real everywhere; it’s just smaller there.

Why this quietly breaks measurement

Most automated visibility tools face a choice: measure the cheap, scriptable, stable API, or do the harder work of driving the consumer interface the way a person does. That choice determines what their numbers mean.

The good news is that the serious dedicated tools have moved to the interface — Ahrefs’ Brand Radar says it runs prompts through the public web interfaces to reflect real user experience, and Profound and Otterly both market monitoring the front-end rather than API calls. The catch is that plenty of cheaper tools and general rank-trackers still hit the API because it’s easy, and some don’t disclose which surface they use at all. When a dashboard reports your “ChatGPT visibility,” the honest question is which ChatGPT — and if the tool won’t say, the safe assumption is the one your customers never see.

What to do about it

You don’t need to become an expert in system prompts and grounding tools. You need one habit: ask any measurement — a tool, an agency, a free scan — which surface it read. The consumer app, logged out, the way a customer sees it? Or an API proxy that’s cheaper to run? If it’s the second, the number is directional at best, and worst on exactly the local, current questions that decide who gets hired.

That’s why Nameworthy measures the consumer interface on purpose, tells you which surface each result came from, and keeps every run on file so you can check the work. It’s slower than a scripted API call. It’s also the only version that matches what happens when your next customer asks. For the ChatGPT-specific version of this, see why the ChatGPT API doesn’t match the app; for the broader case on measurement tools, we pulled a rank tracker apart on the same point.

Common questions

Do AI-visibility tools measure the API or the consumer app?
It varies, and it's the first thing to ask any tool. The serious dedicated tools now measure the consumer interface on purpose — Ahrefs Brand Radar states it runs prompts through the public web interfaces, and Profound and Otterly market that they monitor the front-end rather than the API. But cheaper tools and general rank-trackers often query the API because it's cheap and scriptable, and some don't disclose which surface they use. If a tool won't tell you, assume the API and treat the numbers as directional.
Is the API-vs-app gap the same for every engine?
No. It's largest for ChatGPT, Google's AI Mode, and claude.ai, which each add a hidden system prompt, memory, model routing, and product features on top of the raw model. It's smallest for Perplexity, whose Sonar API was built search-first, so it browses and cites by default much like the app. The gap is real everywhere — it's just narrower on Perplexity.

The report is free

See your visibility measured on the surface customers use.

Six engines, the consumer interface, every run on file — in 48 hours.

Get your AI visibility report