Skip to content

Recognition is not discovery

Typing your own brand into ChatGPT and getting a flattering answer proves the model has heard of you. It says nothing about whether a buyer who has not would ever be told.

Published 2026-08-19 · last reviewed 2026-08-19 · 3 sources

Every founder has done the test. Open ChatGPT, type the name of your own business, read a paragraph that describes you accurately and kindly, and conclude that you are visible in AI search. The measured gap between that test and the one that matters is the widest number in this field. Across 112 startups, ChatGPT recognised 99.4% of the products when they were named in the question, and surfaced them in 3.32% of organic discovery queries, the ones a buyer types when they do not yet know the name. Perplexity fell from 94.3% to 8.29% on the same test. That is one preprint, by one author, on two models, and that limitation belongs on this page as plainly as the number does. But the direction is not in doubt, and it is the reason our frozen question sets never contain the client's name.

Why the flattering answer is guaranteed

A language model asked "is Acme any good?" has been handed the answer's subject. It will retrieve what it can about Acme, or recall what it has, and produce a description. The question contained the brand, so the answer contains the brand; the test has measured whether the model can complete a sentence. It has not measured the only decision that sends a buyer anywhere: when a person who has never heard of Acme asks for the best option in Acme's category, does Acme come up? Those are two different retrieval problems. The first is a lookup. The second is a competition, against every other business the model associates with the category, weighted by whatever the engine has read about each of them. Ninety-nine point four percent on the first and three percent on the second is what it looks like when a business has a name but not a reputation the engines have learned.

One screenshot is not a measurement

The second reason the home test misleads is variance. SparkToro put twelve prompts through three AI tools via 600 volunteers, 2,961 runs in total, and found less than a one in a hundred chance that ChatGPT or Google's AI would return the same list of brands in any two of a hundred responses. The answer you screenshot is one draw from a distribution. The next draw names different brands, in a different order, and may not name yours. The survey of the controlled literature puts seven to eight repetitions per prompt as a reasonable starting point for a stable estimate, and notes that in the configuration studied, 57.8% of ChatGPT repetitions did not even activate web search, so some of those draws came from model memory and some from the live web. A single answer tells you which kind you happened to get. It cannot tell you the rate.

What a real test looks like

Write the questions a buyer would type before they know you exist. Use your own Search Console for the ones people already type, and the engines' own answers for the ones they prompt. Leave your name out of every one; a question that contains the brand guarantees a mention and measures nothing. Freeze the set, so the same questions are asked every time and a later improvement cannot be an easier test in disguise. Ask each question repeatedly, on each engine, with web search on, because that is what a buyer sees. Have a different model read each answer and record whether you were named and whether your pages were among the sources, because a model asked whether it mentioned you will say yes either way. Then report the rate with its band, and the gap between the named question and the blind one. That gap is the number the home test hides, and it is usually the first number a founder has ever seen that describes their actual position.

Run the gap on your own business

The form below queues exactly that test for your domain: one family of prompts that names you, one that describes your category without you, eight repetitions of each on every engine we run. Because those are real engine calls with a real cost, a person approves each run rather than a web form firing it, and the result arrives by email with its sample sizes and intervals rather than as a bare percentage. If the sample cannot separate the two rates, the email says exactly that, because a bare percentage the interval does not support is the screenshot problem wearing a lab coat. One run per domain per day, keyed on the registrable domain, so a subdomain is not a second ticket. The questions used are shown in the email, so you can judge whether the blind one is the question your buyers ask; if it is not, tell us and the measured diagnostic starts from the questions you would choose.

RECOGNITION GAP TEST

Two prompt families, eight repetitions each, on ChatGPT, Claude and Perplexity. We queue it, a person approves it (it is real engine spend), and the result is emailed, usually within a business day.

Eight repetitions per prompt on ChatGPT, Claude and Perplexity with web search on; one named family, one blind. One test per domain per day. No card.

SOURCES

  1. Across 112 startups, ChatGPT recognised 99.4% of products when named but surfaced them in only 3.32% of organic discovery queries; Perplexity fell from 94.3% to 8.29% (Sharma 2026, a single-study preprint with two models).

    Martinez, Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026), arXiv:2607.14035 · 2026-07-15 · 112 startups, 2 models

  2. Under a 1 in 100 chance that ChatGPT or Google's AI returns the same list of brands in any two of 100 responses (600 volunteers, 12 prompts, 3 tools, 2,961 runs).

    SparkToro (Rand Fishkin), New research: AIs are highly inconsistent when recommending brands or products · 2026-01-28 · 2,961 runs

  3. Seven to eight repetitions per prompt is a reasonable starting point for a stable estimate; in the configuration studied, 57.8% of ChatGPT repetitions did not activate web search (Schulte et al. 2026).

    Martinez, Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026), arXiv:2607.14035 · 2026-07-15

Income Lab (Pepform Pty Ltd). "Recognition is not discovery". https://incomelab.me/findings/recognition-is-not-discovery. Accessed 2026-09-22. Licensed CC BY 4.0.

READ NEXT