Skip to content

Comparing toolsBuying one of these

How to compare AI visibility tools, and the three questions to ask first

Every tool in this category will tell you where you rank when somebody asks an AI assistant what to buy. The scores are not comparable with each other, and most of them are not comparable with themselves from one week to the next, because an assistant does not give the same answer twice. Three questions separate a measurement from a number.

Does the same question give the same answer twice?

No. Ask an assistant the same buying question twice and it will often name a different set of brands in a different order, so any single reading is one sample of something that moves.

This is the fact the whole category is built on top of and the one it is quietest about. A large language model does not look up an answer, it generates one, and two generations from the same prompt can name different brands. So a score built from one pass at a question set is one sample, and the honest question about any score is how far it would move if the same measurement were taken again.

That is what a margin of error is for. Of the 7 vendors whose public pages could be read on 5 September 2026, none published one, and none published any acknowledgement that the same prompt gives different answers.

So we measured it on our own question set and published the result. Across 1,017 cases where one assistant was asked one question more than once inside a single measurement, it named an identical set of companies 21.5% of the time, and the five assistants agreed on a single top recommendation for 11 of 166 questions. The study below this page carries the method and every number.

It matters commercially rather than academically. Without a margin of error there is no way to tell a real improvement from the measurement moving on its own, so a team can spend a quarter on work that changed nothing and read a rise, or do the right work and read a fall.

How many prompts is the score built from?

Ask for the number. A score built from a handful of questions moves several points on its own; a score built from hundreds moves less. Of the vendors we could read, none published the figure.

Every score in this category is a proportion: how often you were named, out of how many chances. The bottom half of that fraction decides how much the top half can be trusted, and it is the number nobody prints.

It is not a hard number to publish. It is the count of questions in the set, multiplied by the assistants asked and by how many times each was asked. Of the 7 vendors whose public pages could be read, none of them stated it.

What does the tool do when it cannot measure something?

This is the question that separates the tools, because the answer is invisible in a demo. A tool that reports a zero where it measured nothing will tell you a step of your funnel has collapsed when in fact nobody looked at it.

Measurement fails constantly and for ordinary reasons. An assistant times out. A provider refuses a request. A page cannot be fetched. A brand publishes no price, so there is nothing to compare a quoted price against.

There are two ways to handle that and only one of them is honest. Report the gap and score what was measured, or fill the gap with a zero. A zero reads as a finding about your brand and is a fact about the instrument, and it is worse than a gap because it is actionable: somebody will go and fix a problem that was never there.

Who else is in this category?

Eight products sell measurement of brand position in AI assistant answers. Naming them is what makes the counts on this page checkable rather than a claim about an unnamed market.

Profound, Peec AI, Scrunch AI, Evertune, Otterly, AthenaHQ, Rankscale and Brandlight. Ahrefs and Semrush both ship an AI visibility feature inside a larger search suite, which is a different kind of product and a different comparison.

On 5 September 2026 the public pages of all eight were fetched and rendered. 7 served their content. Otterly refused our crawler on every page, so nothing on this page counts against it either way: an absence we could not look for is not an absence.

Six of the eight state which assistants they measure, which is the one thing this category does publish clearly and is worth checking against the assistants your own buyers use. The counts above are what a re-run of the same check would produce, and the check is a script rather than a reading, so this page carries the date it was taken rather than a claim about today.

What does Delphi Monitor publish?

The margin of error travels with every score, the number of answers a run is built from is printed on the page, and a measurement that could not be taken is reported as a gap rather than scored as a zero.

A run states how many answers it collected and how many of them it managed to read. Every score carries the range the measurement supports, and a change smaller than that range is reported as not yet a result rather than as an improvement. Where a step could not be measured, its weight is redistributed across the steps that could, and the page says which one was missing and why.

None of that is a feature so much as the consequence of one decision: a number that cannot be checked is not worth showing. It is also why the counts on this page are a script anybody here can re-run rather than a claim from a sales deck.

Which of the five is holding your brand back?

Three questions, and the answers decide whether you are buying a measurement or a number. We publish ours because a figure nobody can check is not worth showing.