Skip to content

How the number is made

Most AI visibility reporting is one good answer, framed. This page sets out, in full, how we measure whether ChatGPT, Claude and Perplexity name or cite a business, how we report uncertainty, and when we will tell you a month did not move.

What we measure, and what we do not

Three assistants, asked the questions buyers ask

We measure what three AI assistants say when a buyer asks a question in a client’s category. The assistants are ChatGPT, Claude and Perplexity. Each question is sent to each of them through their APIs with web search enabled, and every answer is stored.

A separate model reads each answer and records whether the business was named, whether its site was returned as a source, and which other businesses appeared. Nothing is installed on the client’s site and no access is required for the measurement itself.

Two presences, and three pillars left unscored

Answer presence is the business named in the answer text. Citation presence is the business’s own domain among the sources the answer cited, which can be measured only when the engine searched the web. Share of voice is measured from the same answers.

The report has three further pillars that come from the website audit. They stay audit only and report no value. A guessed number would look the same on the page as a measured one.

The rate describes the questions we asked

Every rate we publish is movement on a frozen set of questions, across a stated list of engines. It makes no claim about questions we did not ask, or about the category as a whole. That choice decides how the interval is built, and it is set out under the rate below.

A number with no stated scope gets read as a verdict on everything. Stating the scope is what lets the interval be exact about the thing it describes.

The questions

Observed before they are chosen

Candidate questions are ranked from what people have been seen asking. Search Console queries, which people typed, carry the most weight. Buyer questions found in the engines’ own stored answers come next, then questions a monitoring vendor suggests, then questions we wrote ourselves. A question we wrote cannot outrank one that was observed.

Each candidate is ranked on demand, on whether it leads to a page that makes money, and on headroom: a question where the site ranks in Google but is not cited in AI answers has the most room to move, and one where it is cited everywhere has none. Questions about discounts and promo codes are cost questions and are kept, because for an affiliate site they are often the ones that convert best.

Our own comparison site’s first frozen set was written from the category, and none of its ten questions matched anything a person had searched for. The query that earned 11 of its 14 organic clicks was on no scoreboard for a month.

We never let the engine see your name

Asking an engine (an AI assistant we measure) “is [your brand] good?” guarantees a mention and measures nothing at all. We use the questions a real buyer types, such as “what’s the best X for Y?”, and never name the client in any of them. Then we measure whether you surface unprompted.

Any generated question that names the brand or its domain is discarded before the scan runs. If you show up, you earned it. Questions containing the client’s own name sit in the taxonomy (our classification of question kinds) and are never measured.

An answer to “is this business any good” names them whatever the engines think of them, so it measures the question instead of the business. Every question we ask is one a customer asks without knowing the client exists. See brand-blind.

Every question states how it makes money

A question is frozen only when it names its route to revenue, for example an affiliate click to a named merchant from a named page. A route recorded for a page that does not exist yet still counts as money. The route is stored beside the set and kept outside the fingerprint, so older sets still verify.

Freezing is permanent. A question with no route to revenue would be measured forever and would drag the mean for every question that can pay. We once recommended building a comparison page for a category the client held no affiliate in, and it would have won citations worth nothing.

Frozen by a person, and fingerprinted

Only a set a person approves is frozen, and never as a side effect of an unattended scan. Profile answers we research on a client’s behalf are drafts. The freeze refuses until the client has confirmed their own intake. A category the client states is used word for word; a category we infer leaves the scan exploratory and unfrozen, because the same site read twice produced two different categories.

Once approved, the set is frozen: stored with a hash, a fingerprint computed from the exact questions, and printed on every report. Every later scan of that domain reuses it. A stored set that no longer matches its hash stops the scan. Two scans with different hashes are never compared, by code.

The before and after is only fair when both scans asked the same questions. Regenerate them and an improvement cannot be told apart from an easier test. Freezing is also what makes each question its own control.

New ground gets a second set

Coverage grows by freezing a second set, never by editing the first. A later set refuses any question already frozen in an earlier one, its measurement window starts the day it is frozen, and /proof publishes the baseline set only. A later set can stand down from daily collection to save spend and resume afterwards, still frozen. The baseline set cannot stand down.

Two kinds of question, reported apart

Questions divide into two kinds, and the divide matters. Ask an assistant for a category in a suburb and it answers from Google Maps and review data, so an established business usually wins. Ask it what something costs, how you compare, whether you can help with a specific problem, or who supplies an industry, and it answers from website content and third-party pages, where far fewer businesses have built anything.

A single found rate (the share of answers you appear in) averages those two halves and describes neither. So wherever each half holds at least three questions, the rate is split by kind and each half carries its own band. The classification is applied to the stored questions and kept out of them, which means it re-reads every measurement we have ever taken instead of stranding the old ones.

The same business can read as thriving or invisible depending only on which kinds of question ended up on its scoreboard.

How answers are collected

With web search on

The consumer assistants buyers use search the web before answering a buying question; that is true of ChatGPT, which we measure, and equally of assistants we do not measure. An API call without that search measures the bare model’s memory, a different artefact from the answer a buyer is shown. So every engine in our measurement set is probed with its provider’s own web-search tool: native for Perplexity, and the official search tools for ChatGPT and Claude.

Grounded probing (the engine answering with its web search on) also exposes the sources behind each answer, which is how we can measure citation presence as well as answer presence. Every score records its mode. If the mode ever changes, the system declares a new baseline on the same frozen questions and never compares across the break.

We measure the answer buyers see. The model’s memory is a different test.

Many passes, because one proves nothing

AI answers change substantially from run to run. Ask the same question twice and you can get two different vendor lists. A single pass is uninterpretable: you cannot tell a real result from a coin flip.

So every question is asked again on every collection day, on every engine in our stated measurement set. Each run records whether you appeared. The rate across those runs is the measurement.

The engine profile is part of the instrument

How much an engine may do per answer is set by an engine profile: how many web searches Claude may run, and how much reasoning ChatGPT’s model uses. The profile is recorded on every daily row and every score. A measurement window that mixes profiles is refused, and a change of profile starts a new baseline.

From 15 to 29 August 2026 our collection setting allowed Claude one web search per answer. The setting was restored to two searches per answer on 12 September 2026, and figures either side of that change are not compared.

An engine allowed fewer searches can find fewer sources. A rate that moved because the setting moved would be credited to the client’s work.

One collection day, dated in Melbourne

The collector stamps every row with a Melbourne date. A server’s default date is UTC, ten hours behind Melbourne, and a run either side of 10am could split one pass across two dates and count it twice in a window. One pass per day means one pass per day where the buyers live.

Checked before any spend, saved as it goes

Before a pass spends anything, every engine and the grader are called with a time limit on each call. If one is out of credit the day is refused and the operator is told which provider to top up. A graded answer needs a grader, so a grader that runs dry stops the pass.

Each question’s answers are saved as soon as they are graded. A pass that is killed keeps what it collected, and a re-run fills only the missing cells. The same cell cannot be collected twice in one day.

On the first real collection day the grader ran out of credit at question three and every engine’s cell failed with it. Checking the engines without the grader missed the part that failed.

The engine that answers never scores itself

A separate analyst reads every answer

A model asked “did you mention Acme?” will tell you yes whether or not it did. So the model that produces the answer never grades it. A separate analyst model reads every response and extracts which brands were named, in what order, and whether they were recommended or only listed. A check fails the build if any engine shares the analyst’s model.

Leading questions and a model marking its own homework are the two easiest ways to manufacture a good score. Brand-blind questions close the first; a different model doing the grading closes the second.

The gap in that independence today

Today the analyst is a small Anthropic model. It never grades its own answers, but it does read Claude’s, which is one lab reading its own house. That is less independence than the phrase “independent analyst” promises, and our checks warn about it on every run.

Every score names its grader, and old answers can be read again

Every score records which model graded it. When two scans were graded by different models, the comparison says so, because a different grader can read the same answer differently and part of any movement would be the instrument.

The answers are stored, so when the grader changes we re-grade the stored answers and both ends of a comparison share one reader. A re-grade reads mentions again. Citations were established from the source list the engine returned at collection time, which is not in the answer text, so the stored citation flags are carried through unchanged. A re-grade is saved only on instruction and is refused when more than 10% of rows fail.

A first re-grade that re-derived citations from the answer text turned a real 29% citation rate into 9%. The source list is the evidence, and the text cannot stand in for it.

What counts, and what leaves the count

A failed probe is missing data

A probe that failed, from a rate limit or a timeout, leaves the denominator (the count of answers a rate is divided by). It is never counted as a miss.

A rate-limited engine has told us nothing about whether it would name you. Counting its silence against you would report an outage as invisibility.

Claude’s figures use completed searches only

From 15 to 29 August 2026 our collection setting allowed Claude one web search per answer. 166 of its 261 answers in that window say the search stopped early (“I’ve hit the search limit for this turn”), and each of those still carried the source list from the search that did run.

An answer that reports its search stopping did not use those sources, so counting it would state a rate for evidence the answer never had. Those answers are excluded from every Claude figure we publish, the way a failed probe is, instead of being carried with a caveat. The rows stay in the database exactly as collected; the exclusion is applied wherever a rate is computed. The rule is tested on real sentences, including ones that mention a limit without describing a stopped search.

A truncated answer is not a measurement of what the engine would find.

Named and cited

Two rates that can sit far apart

The named rate (answer presence) counts answers whose text names the business. The cited rate (citation presence) counts answers whose sources include the business’s own domain, read from the source list the engine returned. For a comparison or affiliate site, being the source is often the win, because the brands named in the answer are the products being compared.

On our own comparison site, measured through a retired third-party feed on 8 and 10 August 2026 (87 and 147 answers), the site was cited as a source in 29% and 31% of answers and the brand was named in 1.1% and 0.7%. On our own daily collection from 15 to 29 August 2026, 624 completed graded answers on 38 questions, the site was cited in 136 (22%), across 21 of its pages, and the brand was named in 12.

These are snapshots of one property; we hold no series that shows which rate moves first.

A report that counts only one of the two can call a working source invisible.

The Visibility Score, in cases you can check

A weighted blend of rate and position

Visibility Score is a weighted blend of how often and how prominently the business appears, because “named last, grudgingly” and “recommended first” are different business outcomes, and a mention rate prices them identically. The cases below are asserted by an automated check on every push.

CaseScore
Never mentioned in any answer0
Always named first, and recommended100
Named #1 vs named #5, same mention rate85 vs 63
Engine probe failed (rate limit, timeout)excluded
Every probe failed0

The rate and its band

The headline rate

The headline rate is the mean of the per-question rates, one rate for each question on each engine. It is kept apart from a pooled rate (all hits over all answers), which drifts from the mean when probes fail unevenly. A fortnight of daily passes on twenty-five questions across three engines is about 1,050 answers, which supports a 90% band of a few points either side.

The report shows the band only when every question has more than one graded answer on each engine. That is a number worth acting on.

No interval on a single answer

An interval needs more than one sample. A question asked once returns no interval at all. The report checks the data itself before it shows one, separately from whether a plan includes intervals, and both checks stay separate on purpose.

An interval drawn around one answer would present a guess as a range.

A single question shows counts

A single question on a single engine is only fourteen answers, giving a 90% interval of about 20 points either side. That cannot support a decision. So per-question rates are shown as directional counts (cited 9 of 14) with no band. The code sets 20 answers as the least a per-question band would need, and no page shows one.

The headline band, z = 1.645

rate = (p₁ + … + pₖ) / k
se   = √( Σ pᵢ(1−pᵢ)/nᵢ ) / k
band = rate ± z·se

Here k is the number of question-and-engine rates, pᵢ one of those rates, nᵢ the answers behind it, and z = 1.645 sets the band at 90%. The band is centred on the rate we publish.

Wilson interval, the fallback at no hits or no misses

centre = (p̂ + z²/2n) / (1 + z²/n)
half   = z·√( p̂(1−p̂)/n + z²/4n² )
         / (1 + z²/n)

When there are no hits, or no misses, each pᵢ(1−pᵢ) is zero and the band above would shrink to a point. There we use Wilson over all the answers, where p̂ is the overall rate and n the total answers. Wilson is used in place of the normal approximation (the textbook formula), because at p̂ = 0 the normal interval collapses to zero width and claims certainty we do not have. A rate of zero is a real and common reading on one half of a scoreboard.

Why the band carries no correction for clustering

The thing we estimate is the frozen set on the stated engines. For a mean over that fixed grid of questions and engines, each cell’s own sampling error is the only error there is. A cluster-robust correction would widen the band to cover questions we might have asked, and we make no claim about those.

A wider band would look more cautious while answering a question the report never asks, and it would hide real movement on the questions it does ask.

The before and after test

Paired on the question

A change counts as real on the report only when the same questions, asked again, move by more than their own noise. Each question’s after is compared with its own before, and the 90% interval on the question-by-question difference must exclude zero. Each frozen question is its own control. The rule is fixed before we look.

If a month does not clear that bar, we will tell you it was within noise, and why.

Why overlapping bands are the wrong test

Checking whether the before band and the after band overlap throws away the pairing, and it is stricter than asking whether the difference is non-zero. It would report real movement as “cannot tell” and deny a client credit for work they paid for.

Pairing needs each question’s own rate stored on the score. Older scans are filled in from their stored answers. Where that cannot be done, the comparison falls back to the overlap test and says so in its output.

Controls

Your movement, minus the drift

Engines change on their own. A model update can lift or sink your visibility in a week with nothing you did. If we only measured you, we would take credit for the weather.

So a control is measured on the same frozen questions, in the same answers, at the same moment: nothing extra to run, and no chance of the control being measured on a different day or a different test. The attributable lift is the difference in differences, computed from the stored scores; it is a separate test from the report and is not yet printed on it.

lift = (client after − client before)
     − (control after − control before)

Your change on each question minus the control’s change on that question, averaged, with its own 90% interval. Whatever moved the whole category moves the control too, and subtracts out. What is left is attributable.

Matched to what the business needs to win

For a brand that needs to be named, the control is the rivals named at scoping. For a comparison or affiliate site, where the brands in an answer are the products being compared, it is every other domain used as a source in the same answers; that is the control committed for our own comparison site on 10 August 2026.

On our own comparison site, the five rival brands named as controls appear in zero of 90 answers. A control built from them would subtract nothing and present the raw change as attributable.

Chosen before the result

The control is committed in writing, with a timestamp, before the outcome measurement, and it cannot be overwritten. A control committed later is reported as not pre-specified. The test refuses to run across any change of question set, source, web search mode or engine list.

Rival names suggested by a monitoring vendor are treated as signal only. Their mapping from a brand to a website is never used as a control, because it has mapped rival names to a pen manufacturer and a networking company.

A control picked after seeing who moved can be made to show anything.

Measurement windows

Consecutive, and never overlapping

Each window starts the day after the last saved measurement of the same question set. The first starts on the day the set was frozen, or on a stated start date. A window is saved once, when it holds enough collection days.

An earlier version saved a measurement on every trigger once fourteen days existed, so each window contained the last. An after that contains its before cannot show a change.

Refused when too little was graded

A measurement is refused when more than 10% of its rows could not be graded. The same refusal applies to a window that mixes engine profiles.

Failed rows leave the denominator, so a half-graded window would publish a confident rate computed on whichever answers came back before the grader stopped. That is a biased sample presented as a clean one.

When the instrument changes

The engines are not a fixed ruler

Labs ship new models, retraining changes what a model already believes, and retrieval changes which pages get read. A number measured today describes today’s engines. When a lab ships a new model, the ground moves under everyone in your category at once: you, and every competitor, on the same day.

That is the reason a single measurement is a photograph and the series is the evidence. Your competitors are measured in the same answers at the same moment, so when the ground moves, it moves under them too; a one-off report ages with the engines, and the series does not.

What every score records

A comparison is made only between scores that share every one of these. When a series crosses one, we compare within the newest run of scores that share them all, and state what was excluded and why.

  • The question setRecorded as its hash. Different hashes are never compared.
  • The measurement sourceOur own engine calls and rows ingested from a third-party feed are separate modes, recorded on every score.
  • Web search on or offGrounded and ungrounded runs are different tests.
  • The engine listAdding or removing an engine changes what the mean is a mean of.
  • The engine profileHow many searches an engine may run changes what it can find.
  • The graderRecorded on every score. A change is flagged, and stored answers can be re-graded so both ends share one reader.

A client who crossed a boundary once still deserves an answer about everything since. Refusing all comparison left one comparable pair of measurements, two days apart, reported as not comparable.

What we will not claim

  • Movement that did not clear the paired test. A month inside the noise is reported as within noise, with the reason.
  • Anything about questions we did not ask. The rate describes the frozen set on the stated engines, and nothing wider.
  • An engine we did not query. Every report states the AI assistants it covers.
  • A comparison across a change of question set, measurement source, web search mode, engine list or engine profile.
  • A value for the three audit pillars we do not measure. They report no value instead of a guessed one.
  • A Claude rate that counts answers whose search stopped early.
  • A monitoring vendor’s dashboard score as a client result. Those scores are read as direction only.
  • Visits from AI assistants as an exact count. Consent banners, stripped referrers and AI Overviews traffic arriving as ordinary Google all hide some, so the count is a floor.
  • What anyone says to an AI assistant in private.

What we can measure today, and what we cannot

What we can measure

What ChatGPT, Claude and Perplexity say
when your customers ask, with web search on. Every answer is stored.
Whether you were recommended by name
and separately whether your pages were the sources behind the answer. Most tools count only one.
Who gets recommended instead of you
on the same questions, measured at the same moment.
Before and after
the same questions, compared question by question. The headline rate carries a 90% band; a single question shows counts. A change is flat when the 90% interval on the paired difference includes zero.
Visits arriving from AI assistants
in your own analytics, read with your permission, reported as a floor (the real number can only be higher).

What we cannot, yet

We have not yet published a clinic
Our published measurements are our own properties. The first third-party results come from the design-partner round, and they will be published whether or not they move.
The free website check
reads your website only. It does not look at your Google Business Profile, your listings, or anything said about you anywhere else.
Google’s AI Overviews
is not something our own rig (our own engine calls) probes today, nor are other assistants beyond the three. Our measurements of 8 and 10 August 2026 came through a retired third-party feed that did include it, and those rows are labelled wherever they appear. Our 21 and 22 July 2026 measurements were our own engine calls without web search.
We cannot see private conversations
anyone has with an AI assistant. Nobody can. Anyone claiming to measure them is guessing.
We cannot make visibility move in three weeks
Engines re-read the web on their own schedule, and the chain from a change to a moved rate has six steps. Real movement shows in about ninety days, and we publish our own flat months to prove we will report yours.
We cannot hold the engines still
Labs ship new models and retrain old ones, and the answers change for your whole category at once.

Ninety days is arithmetic

A change you make today cannot appear in an answer today. Six things have to happen first, the engines run on different clocks, and the citation you earn decays whether or not anyone acts on it. All of that is set out with its evidence on how pages become sources. What decides which sources an engine reaches for in the first place differs so much between engines that only 2.7% of cited domains are shared across five of them, which is set out in the research we build on.

Our own numbers

The same standard, run on ourselves

We run this exact standard on our own business. At our first measurement the AI engines named us in zero of 149 answers. That run is published on /data; it was an ungrounded run (engines answering without web search), so the grounded measurements that follow, on /proof, start a new baseline beside it, as they move or fail to.

Our own numbers, including the unflattering ones, are on Proof. Our live measurement ↗

Corrections

  1. Every Claude figure on this site is re-scoped to its completed searches. 166 of Claude’s 261 answers from 15 to 29 August 2026 report their search stopping early while still carrying a source list, and counting them inflated every Claude rate. The corpus behind our published figures is 624 completed answers, not 790.

  2. This page said flat months auto-draft as flat. No such generator exists in the code, so the sentence is removed.

  3. This page said numbers before 12 August 2026 came through a retired third-party feed. Only our 8 and 10 August 2026 measurements did; our 21 and 22 July 2026 measurements were our own engine calls without web search.

  4. This page said the verdict is computed from the bands, and set the difference-in-differences against the control group’s band. Each test is the 90% interval on a question-by-question difference.

  5. This page said every rate carries a 90% band, that movement inside it is flat, and that every published rate is split by kind. The headline rate carries a band once every question has more than one graded answer on each engine, a single question shows counts, the split needs three questions per half, and flat means the paired interval includes zero.

  6. This page showed the Wilson formula as how the band is made, under a heading calling the headline rate pooled. The band is a variance sum over the per-question rates, centred on their mean, the rate we publish; Wilson is the fallback at no hits or no misses.

  7. This page said the control is the rival brands named in your category. That holds for a brand that needs to be named; for a comparison or affiliate site, including our own (committed 10 August 2026), it is every other domain used as a source in the same answers.

  8. This page said the cited rate moves before the named rate; we hold no series showing that. Its three answers in ten came from a retired feed on 8 and 10 August 2026 (87 and 147 answers); our own 790 answers from 15 to 29 August 2026 show the site used as a source in 162 (21%) and the brand named in 12.