AnswersSample size and confidence
How many questions does an AI visibility measurement need?
There is no single right number of questions. What decides a measurement is how many answers sit behind each figure, which is the question set multiplied by the assistants asked and by the number of times each is asked, and how wide a confidence band you are willing to accept around the result. The honest form of the answer is a range, and a figure published without one is asking to be taken on trust.
Why does the same question get a different answer each time?
AI assistants do not answer identically twice. Put the same buying question to one assistant more than once and the shortlist can come back in a different order, with different names on it and different reasons attached.
This is a property of how the assistants work rather than a fault in any one of them. A model samples from a range of possible replies, so two runs of the same prompt take two different paths through it. Where an assistant reads the live web as it answers, the pages it read this morning are not the pages it read last week, and the answer moves with them.
The consequence for measurement is the whole of this page. A single answer is one draw from a distribution. It is evidence that the answer is possible. It is not evidence of what a buyer is usually told.
What is a confidence band, and why does every figure carry one?
A confidence band is the range a figure could reasonably sit in, given how far the answers behind it disagreed with each other. Delphi computes one from that disagreement and publishes it wherever the Delphi Index appears, so the score reads as a range rather than as a point.
The band is the spread between the assistants on the same questions, taken as a standard error and scaled to a ninety five percent interval. A category where every assistant says roughly the same thing produces a narrow band. A category where one assistant puts you first and another leaves you out entirely produces a wide one, and that width is a finding rather than a defect in the run.
Two properties of it are worth knowing. It is floored at one point, so the product can never report absolute certainty about a figure derived from answers that vary. And where fewer than two assistants answered there is no spread to compute at all, so no band is published and the run states that rather than printing a zero.
Why is one run of one assistant not a measurement?
One run of one assistant is a single draw from a system that answers differently each time. A measurement puts the same question set to five assistants, more than once each, and keeps the disagreement between those answers rather than averaging it away.
Three things have to hold before a set of answers is worth calling a reading.
- Repeats. Every question goes to each assistant more than once on a full run, so one unusual answer cannot set the score by itself.
- Breadth. Five assistants answer the same set: ChatGPT, Gemini, Claude, Perplexity and Grok. A brand can be the obvious recommendation in one of them and missing from another.
- The disagreement is kept. It becomes the band around the figure instead of being smoothed into an average, because how far the assistants differ is the part a buyer actually experiences.
An assistant that did not answer is disclosed rather than filled in. The Index is computed from the assistants that did answer, the shortfall is stated beside it, and an answer from an earlier run is never reused as a fresh one.
How many questions is enough?
How many questions is enough depends on how wide a confidence band you are willing to accept around the result. The answers behind a figure are the questions multiplied by the assistants asked and by the repeats, and there is no universal minimum that holds across every category.
That is the shape of the arithmetic and it is as far as it can honestly be taken. What number of questions gives you a band you can act on depends on things that differ by brand: how far the assistants disagree in your category, how many of your questions carry buying intent rather than research intent, and how small a monthly movement you need to be able to call.
Two limits here are fixed. A question set is capped at forty questions and a run at seven hundred answers, because every question is put to every assistant more than once and the size of the set is what a run costs. Inside those, the set is yours to edit, and every run is pinned to a version of it, so one run compares with the next on the same questions.
The figure worth asking any tool for is not the question count. It is the number of answers behind the score, and the band that number produced.
When is a change a result, and when is it noise?
A change smaller than the confidence band is reported as no confirmed change rather than as a movement. Delphi applies that test through one shared rule, so two screens can never disagree about whether the same change was real.
The test is deliberately blunt. If the size of the change is larger than the band it is reported as a confirmed change. If it is not, the reader is told that the figures moved by less than the margin of error of the measurement itself, and that it is too early to call any of it a result.
There is a third answer and it is not no. Where the band could not be computed, because too few assistants answered, the product says it could not tell whether the figure moved. Could not tell and did not move are different statements, and collapsing them lets a product report a result it never measured.
Two further cases are refused rather than reported. Where the way answers are read or scored changed between two runs, the step between them is not published as a movement, because it was made by the method rather than by the brand. And three of the five steps carry no margin of their own: Seen and Chosen each have an observable that can be split by assistant, while Found, Trusted and Bought do not, so a change on those three is never called a confirmed result, and the run says it could not tell rather than reporting a verdict it did not measure.
What can this method not see?
An AI visibility measurement is a designed sample of buyer questions on the day it ran, not a census of what everybody asked. It cannot reproduce a personalised answer, it does not read retailer owned assistants, and it says nothing about what happens after a buyer leaves the assistant.
These are properties of the method rather than gaps waiting to be filled, and they are the reason the rest of the page is worth anything.
- A sample, not a census. The reading covers the questions in your set on the day the run happened. A question nobody put in the set was not measured.
- Personalisation and memory. An answer shaped by the history of one person is not reproducible, and no attempt is made to model it. What is measured is the answer a new buyer with no history receives.
- The assistants change underneath the measurement. A model update can move a score with nothing changed on your side. Every answer records the model that produced it, so a movement caused by a model change can be told apart from your own progress.
- Retailer owned assistants are not read. Nothing here measures what an assistant inside a retailer tells a shopper, and no figure implies otherwise.
- No commercial outcome is claimed. A position in AI answers is never converted into revenue anywhere in the product.
How should you interrogate an AI visibility number?
Four questions establish whether a visibility number is a measurement or a figure: how many answers is it computed from, what is the range around it, what happens when the same question is asked twice, and what is reported when a movement is smaller than that range.
Take them to any tool, this one included. Each has an answer a vendor can give in a sentence, and an evasion that is easy to recognise once you are listening for it.
- How many answers is this computed from? Questions, assistants and repeats, as three numbers rather than one.
- What is the range around it? A figure with no range has not been checked against how much the answers behind it vary.
- What happens if you ask the same question again? If the answer is that nothing changes, either the sampling is very deep or it happened once.
- What do you report when a movement is smaller than the range? The honest answer is no confirmed change, and it is the one nobody enjoys giving.
Our own answers are on this page and in the method, in the same words. A confidence band is published beside the Index, a movement smaller than that band is reported as no confirmed change, and where the band could not be computed the product says so instead of quietly reporting stability.
Related answers
Which of the five is holding your brand back?
A measurement you can argue with is worth more than a number you cannot.