Skip to content

ResearchMeasured, not argued

Do AI assistants agree on what to recommend?

Almost never, and they do not reliably agree with themselves either. We put 188 questions to 5 assistants across 8 brands at 4 companies, and kept every answer, 3,116 of them. Asked the same question more than once inside one measurement, an assistant returned the same set of names 21.5% of the time. On the 166 questions where at least three assistants named a company, all of them landed on the same top recommendation 11 times. Here is the whole method and every number.

Does the same assistant give the same answer twice?

Usually not. Across 1,017 cases where one assistant was asked one question more than once inside a single measurement, it named an identical set of companies 219 times, or 21.5%. In the other 798 the list of names changed between one asking and the next.

This is the finding that governs every other number in the category. A language model does not look up an answer, it generates one, so two askings of one question are two samples of something that moves rather than two readings of a fixed fact. The measurement here is direct: one question, one assistant, asked more than once inside the same measurement, comparing the set of companies named each time.

Asking within a single measurement is what makes it a fair test. Pooling answers from measurements taken months apart would count a market that genuinely changed as an assistant failing to repeat itself, which is a different and much weaker claim.

The practical reading is that a single pass at a question set tells you what one assistant said once, and the useful question about any visibility score is how far it would move if the same measurement were taken again.

Do the five assistants agree with each other?

Rarely. On the 166 questions where at least three assistants named a company, the most-named company was the same across all of them on 11 questions. On the other 155 they disagreed, with 3.34 different top picks per question on average.

For each question we took the company each assistant named most often, then asked whether the assistants had landed on the same one. The average question produced 3.34 different top picks between the assistants that answered it.

So there is no single ranking to be at the top of. There are five, they move, and the company a buyer meets depends on which assistant that buyer happens to open.

Does it matter which assistant you check?

Substantially. Of the 2,758 companies named anywhere in the study, 2,088 were named by exactly one of the 5 assistants, which is 75.7%.

Three companies in four appeared in front of one assistant and none of the others. Checking one assistant is a real reading of that assistant and a poor proxy for the rest.

How many companies each assistant will name at all is closer than the disagreement above suggests, at a spread of 1.4 between the widest and the narrowest. ChatGPT named 880 distinct companies across these questions and Perplexity named 645. So they name similar numbers of companies and largely different ones, which is the more useful way round to state it.

  • ChatGPT: 880 distinct companies named
  • Claude: 857 distinct companies named
  • Gemini: 840 distinct companies named
  • Grok: 788 distinct companies named
  • Perplexity: 645 distinct companies named

How many companies are competing for these answers?

More than any category admits. 2,758 distinct companies were named across 12,639 naming events, and 1,588 of them were named exactly once. The single most-named company took 3.4% of all namings and the top three took 8%.

There is no incumbent holding these answers. The most-named company appeared in about one naming in thirty, the top three together in about one in twelve, and 1,588 companies surfaced a single time across the whole study. That long tail is what a generated answer does: it reaches for something plausible, and plausible is a very wide set.

561 of the 3,116 answers, 18%, named no company at all. Those answers explained how to approach the problem instead of recommending anybody, which is worth knowing if you are waiting to be recommended on a question that is not being answered with names.

What does this mean for an AI visibility score?

It means a score needs a denominator and a margin of error to be readable. A number built from one pass at a question set is one sample of something that changed between askings 78.5% of the time in this study, so the honest form of a score states how many prompts it rests on and how far it would move on a re-run.

This is the practical use of the study, and it applies to any tool in this category including ours. Ask how many prompts a score is built from, ask how many samples per assistant, and ask what the same measurement returned last time. Those three answers turn a number into a measurement.

It is also why repeat sampling is the design rather than a feature. Every measurement here asked each assistant each question more than once, which is what made the 21.5% figure computable at all. A measurement that samples once cannot report its own stability, because it has nothing to compare with.

What this study covers, and what it does not

It covers 3,116 labelled answers to 188 questions from 5 assistants, across 8 brands at 4 companies. It supports claims about how much these assistants varied on these questions, and not about any company's standing today.

Three limits are worth stating plainly, because a study that hides them is a brochure. The panel is 4 companies, which clears our reporting floor and is still small. The questions are the ones those businesses chose to measure, drafted for their own categories rather than sampled from search demand. And these systems change, so a figure here is a reading of a period rather than a constant.

The base is also narrower than the raw measurements. Answers reach these counts once the labelling stage has read them, and unlabelled answers are excluded rather than counted as naming nobody. That distinction decides the headline: an unlabelled answer has no list recorded against it, so counting it as an answer that named nobody would make a repeat asking look like a changed answer.

One more thing is true and belongs here. Our own product was named in none of the 357 answers measured for it. The study is the market we measured, and it did not include us.

Which of the five is holding your brand back?

Every number here came from asking the same questions more than once and keeping what came back. That is the whole method, and it is available to anybody willing to ask twice.