AnswersAI visibility scores
What an AI visibility score is, and why two tools give one brand two numbers
An AI visibility score puts one figure on where a brand stands when buyers ask an AI assistant what to buy. It is a position in answers: not traffic, not revenue, and not a forecast. Two tools measuring the same brand on the same day will disagree, because a score is a reading of a question set put to a set of assistants and counted by one set of rules, and no two tools make those choices the same way.
What does an AI visibility score actually measure?
An AI visibility score measures where a brand stands in the answers AI assistants give buyers: whether it can be read at all, whether it appears, whether it is described accurately, whether it is recommended, and whether it is ready to be bought from. It is a position in answers, not a count of visits and not a share of revenue.
The main unit of measurement is an answer. Delphi puts a brand's real buying questions to ChatGPT, Gemini, Claude, Perplexity and Grok, stores every answer word for word, and reads those stored answers for what they say about the brand. The exact model that produced each answer is recorded beside it, so a figure drawn from answers traces back to sentences a person can read. The first and last of the five also read the brand's own pages, because whether a machine can read a site, and whether the brand is ready to be bought from, are not questions an answer settles on its own.
So the score is a position rather than a volume. It says where a brand stands among the names an assistant offers a buyer. It does not count how many buyers asked. That is a different measurement, and it is not one this figure contains.
How is an AI visibility score calculated?
Delphi builds its score from five sequential gates, Found, Seen, Trusted, Chosen and Bought, each scored out of 100 and combined into a single Delphi Index out of 100, published with a confidence band drawn from how far the assistants disagreed with each other.
The five are in the order a buyer passes through them, and each answers one question:
- Found. Can AI systems access and read the brand at all.
- Seen. Does the brand appear in AI answers for category questions.
- Trusted. Is it described accurately and positively.
- Chosen. Is it recommended at or near the top.
- Bought. Is the brand ready to be transacted when an assistant sends a buyer.
The gates are weighted, and they are not weighted equally. The weights are fixed by the model rather than set by the customer, so two brands measured by Delphi are measured on the same scale, and changing one is a change to the method rather than a setting.
Questions with no purchase intent stay in the set and are excluded from recommendation scoring. A question about how a category works has no recommendation to win, so counting it would create a false zero on Chosen for a brand that failed at nothing.
Why does the weakest gate cap the ones after it?
The gates are sequential, so the weakest one caps every gate after it for a causal reason rather than a scoring one: a buyer who cannot find a brand never reaches the question of whether they trust it. Points sitting behind a low gate pay nothing until that gate moves.
An assistant that cannot read a site cannot describe it accurately, whatever the site says. A brand nobody names cannot be the recommendation, however good its product pages are. So work aimed at a later gate can be finished, correct and worth nothing for as long as the earlier gate is where it is, and a score that averaged the five without saying so would hide the one fact a team needs.
That is why the lowest gate is named as the binding constraint and leads the plan. On a tie the earlier gate in the chain wins, because a buyer meets it first, so fixing it is strictly the right order.
What happens when a gate cannot be measured?
A gate the run could not measure is excluded from the Index and the remaining weights are renormalised, never scored as zero, and the number of gates the Index was computed over is published beside the figure.
Weighting an unmeasured gate at zero states a finding. It says the brand failed at something nobody looked at, and it caps the Index at a ceiling the brand can do nothing about. Excluding the gate and rescaling the rest keeps the number a reading of what was actually measured, and the disclosure is what stops that being a quieter kind of dishonesty: a number computed over a smaller basis is a different number, and hiding that is a lie of omission.
The same rule runs one level down. An Index is computed from the assistants that answered rather than from the assistants that were planned, and where answers were collected and could not be read, that shortfall is stated beside the figures. A gate that was not measured is also never named as the binding constraint, so a plan can never lead with work nobody looked at.
What is a good AI visibility score?
Delphi reads its Index in four bands: Weak below 35, Fair below 55, Good below 75, and Strong at 75 and above. A band is a position rather than a target, and it is only meaningful against scores produced by the same method.
A band word can change without the number meaning anything new. A single point across a threshold moves a brand from Fair to Weak while sitting well inside the measurement's own margin of error, which is why a movement smaller than that margin is stated as unconfirmed rather than plotted as a trend. Wherever the Index appears, the band around it appears too.
A score is comparable only with another score produced the same way. Setting an Index beside a figure from a different tool compares two different question sets put to two different sets of assistants under two different counting rules, and the arithmetic will not reconcile. The comparisons that do hold are a brand against itself over time on the same question set, and a brand against the rivals named in the same answers.
Why do two AI visibility tools give one brand two different scores?
Two tools score one brand differently because they are not measuring the same thing. They asked different questions, asked them a different number of times, put them to a different set of assistants, weighted position inside an answer differently, and drew the line between a mention and a recommendation in a different place.
- Which questions were asked. A score is a reading of one question set. A set weighted toward the corners of a category a brand already owns reads high, and a set built from what buyers actually type reads lower. Neither is wrong and they are not the same measurement.
- How many times each was asked. These models do not answer identically twice. Asked once, a question records one draw. Asked several times, the disagreement between the answers becomes a band around the figure. A tool that asks once has no band, and no way to tell a movement from noise.
- Which assistants were asked. The assistants disagree with each other on the same question, and they disagree about which sources they read. A score over three assistants and a score over five are different measurements even on an identical question set.
- How position inside an answer is weighted. Being the one name an assistant gives, reaching the first three, and being mentioned somewhere in the paragraph are three different commercial outcomes. A tool that treats them alike and a tool that separates them will report different numbers from the same answer.
- What counts as a mention at all. An answer can credit a brand under a sub-brand, a product line or the legal entity. A tool matching the trading name alone scores those as absent. A tool matching loosely credits a brand for a name that was never theirs, which is the same error in the more flattering direction.
The fourth is worth an example. Take a hypothetical set of four brands: Ardent, Kestrel, Trailhead and Nimbus, all invented for this page. An answer that opens by naming Ardent and then warns the buyer away from it puts Ardent in first position. Counted on position alone, that is the strongest possible result. Read for what it tells a buyer, it is the worst answer on the page. Delphi does not count a named-and-warned-against answer as a recommendation, and a tool that counts position alone scores it as a win.
None of the five is a fault. They are design decisions and each one is defensible. What separates a score worth acting on from a number is whether the tool publishes which decisions it made, so the reading can be checked instead of believed.
What does an AI visibility score not tell you?
A visibility score does not say how many buyers saw the answer, how much revenue moved, whether a change you made caused a movement, or what an assistant will say next month. It is a position in a sample of answers, taken on the day it ran.
- Not a count of buyers. Nothing here counts how often a question was asked, so no figure on this page reports demand.
- Not revenue. The score measures a position in answers. It does not measure money that moved because of that position, and nothing converts one into the other.
- Not proof of cause. A gate that rises after work shipped is what happened afterwards. Competitors move, assistants change, and the evidence they read moves with them.
- Not a census. It is the questions in your set, on the day it ran, put to five assistants. It is a designed sample of buyer intent rather than every question anybody ever asked.
- Not stable underneath you. A model update can move a score with nothing changing on your side. Every stored answer records the model that produced it, so a movement of that kind can be seen for what it is.
- Not personalised. What is measured is the answer a buyer with no history receives. An answer shaped by one person's own past use is not reproducible and is not modelled here.
Related answers
- How AI visibility is measuredThe scoring model in full, including what it cannot see.
- How many questions does a measurement need?The question set is the denominator under every figure a score reports.
- Mentioned is not recommendedBeing named in an answer and being the name a buyer acts on are two readings.
Which of the five is holding your brand back?
A score is worth what its denominators are worth, and the ones behind this one are published so you can argue with them.