Skip to content

Methodology

How the number is made, and what it cannot see.

Delphi puts the questions your buyers ask to ChatGPT, Gemini, Claude, Perplexity and Grok, stores every answer word for word, and scores what those answers say about you. This page is the whole method, including the parts that limit it. If you are evaluating Delphi, the last two sections are the ones worth your time.

Five gates, in the order a buyer passes through them

A buyer has to find you before they can see you, see you before they can trust you, trust you before they choose you, and choose you before they can buy. So the weakest gate caps every gate after it. That gate is named the binding constraint and it always leads the plan, however good the others look.

Found
Can AI systems access and read the brand at all.
  • Crawler access. Whether the crawlers behind the assistants are allowed to read the site at all. A refused crawler caps every gate after this one for that assistant.
  • Structured data. Whether pages carry the machine-readable labels that tell an assistant which text is the price, the rating, the ingredient list.
  • Readable claims. Whether key facts exist in the HTML as delivered, rather than appearing only after JavaScript runs or a tab is clicked.
  • Brand recognition. Whether the assistants already hold the brand as a distinct entity rather than an unfamiliar string.
Seen
Does the brand appear in AI answers for category questions.
  • Appearance rate. How often the brand turns up when a buyer asks a category question.
  • Share of the answer. Of all brands named across those answers, how many mentions are the brand's.
  • Breadth. How many distinct question clusters the brand appears in at least once.
  • Category authority. Presence on informational questions. Reported and never scored: a brand absent from a how-does-it-work question has not failed at anything.
Trusted
Is it described accurately and positively.
  • Factual accuracy. Where an assistant and the brand's own published page tell buyers different things. Only scored when at least one claim could be compared.
  • How you are described. Whether the answers speak favourably, neutrally or unfavourably. A named-and-disparaged mention is treated as worse than a neutral one and better than absence.
Chosen
Is it recommended at or near the top.
  • First-choice rate. How often the brand is the recommendation rather than one of several. Purchase-intent questions only.
  • Shortlist rate. How often the brand reaches the first three names, which is the set a buyer actually considers.
Bought
Is the brand ready to be transacted when an assistant sends a buyer.
  • Where to buy. When an assistant recommends the brand, how often it also tells the buyer where to get it. A named retailer counts for more than a marketplace, and a marketplace for more than the brand's own site.
  • How concrete the offer is. How often the answer carries a price, a pack size or a named variant rather than the brand name alone.
  • Obstacles raised. How often an assistant tells a buyer something that stops the purchase: discontinued, out of stock, hard to find, not sold here.
  • Agent checkout readiness. The brand's own site: complete product data, and an interface an AI agent could transact against. Weighted lowest of the four.

The five are fixed and are not a setting. Changing how they combine is a change to the model, and it would be logged and dated here.

The Index, and the band around it

The Index is the weighted sum of the five gates, out of 100. It is a position, not a forecast.

Bands
Weak below 35, Fair below 55, Good below 75, Strong at 75 and above.
Confidence
These models do not answer identically twice. Every question is put to each assistant more than once and the disagreement between those answers becomes a band, published everywhere the Index is published. A movement smaller than the band is not reported as movement.
A gate we could not measure
It is excluded from the Index and the remaining weights are renormalised, and the number of gates scored is stated alongside the number. It is never counted as zero, because a zero says the brand failed at something nobody looked at, and it caps the Index at a ceiling the brand cannot move.
An assistant that did not answer
The same rule. The Index is computed from the assistants that did answer and the shortfall is disclosed. Nothing is ever backfilled from a previous run.

Where a conflict comes from

The finding people ask about most, so it is worth being exact about what it is and is not.

Delphi compares two published things: what an assistant tells a buyer, and what your own page says. Where they differ, both are quoted and neither is declared correct.

That is deliberate. A buyer asking an assistant never sees your page. If an assistant tells them a product is discontinued, the sale does not happen, and it makes no difference to that outcome whether your page is right and the assistant is stale, or your page is out of date and the assistant is reading it correctly. Either way you are losing the sale, and either way there is work to do. Presuming your page is the truth would send half of those cases to fix the wrong end.

Where your page does not address a claim at all, that is recorded as no comparison rather than as agreement. Accuracy is only scored when something was genuinely comparable.

The plan

48 plays. Each carries a trigger evaluated against measured signals, the mechanism it works by, an owner, an effort estimate, and an expected lift stated as a range with its basis.

A play either fires or it does not
The trigger is evaluated in code, not requested of a model. There is no close enough. A signal the run did not measure resolves to false, so a thin run produces a short plan rather than a confident wrong one.
The same inputs give the same plan
The grounded pass has no sampling, no clock and no randomness, and a test asserts it. That is what the Index trend and any quarterly review depend on.
Sequencing
The binding gate leads. A finding that suppresses several gates at once, such as a blocked crawler, outranks even that, because lifting a later gate while the crawler is refused moves nothing. Prerequisites are respected.
Lift is a range
Never a point estimate, never attribution. The basis is printed on the card.

What Delphi does not do

Claim your revenue
Delphi measures your position in AI answers. It does not measure revenue caused by that position and does not convert one into the other.
Read retailer assistants
Delphi does not read Rufus, Sparky or any other retailer-owned assistant, and does not imply otherwise. Retail signals, where present, are correlation and are labelled as such.
Show you a number it did not measure
A widget with no data states what is missing and the one action that fills it. There are no placeholder figures anywhere in the product. The single exception is a clearly labelled illustrative example, which carries a permanent banner saying so and is blocked from every export.
Promise an outcome
No projected figure is presented as a certainty.

The limits worth knowing before you buy

These are the honest edges of the method. They are here because an evaluation that finds them later is worse for both of us than one that finds them now.

A sample, not a census
Delphi measures the questions in your set, on the day it ran. It is a designed sample of buyer intent, not every question ever asked. The set is yours to edit and every run is pinned to a version of it, so a comparison is always like for like.
Assistants change underneath us
A model update can move your score without anything changing on your side. Every answer records the exact model that produced it, so a movement caused by a model change is visible as one rather than mistaken for your own progress.
Personalisation and memory
An answer shaped by one user's own history is not reproducible and Delphi does not attempt to model it. What is measured is the answer a new buyer with no history receives.
Competitors are compared on two gates
Rivals are scored on Seen and Chosen from the same answers as you. They are not scored on Found, Trusted or Bought, because those need their own site audit and claim adjudication. So there is no competitor Index, and the comparison says exactly which two gates it covers rather than implying five.
Bought measures whether the shelf is described correctly
Most brands do not sell on their own website, so auditing it measures a shop that does not exist. Bought reads the shelf the assistants put in front of a buyer: the price they quote, whether they say you can be bought right now, the specifics they assert, and the shop they send people to. Each is compared against what your own pages publish. Your site still counts as the readiness component, weighted lowest, where it explains why an assistant gets you wrong rather than standing in for whether it does.
Delphi never places an order
Nothing here transacts, and no checkout is ever started or completed. Every signal is read from what an assistant says. Where an assistant quotes a price below what you publish, that is not counted against you: retailers discount, and describing the market accurately is not an error. What is counted is the material divergence.
A comparison nobody can make is reported as unmeasured
If your product pages do not publish a price, price accuracy is not scored at all rather than scored as zero, and its weight moves to the parts that could be read. A brand nobody can check is never treated as a brand that got it wrong.
In-store and phone are out of scope
Delphi measures the digital shelf inside AI assistants. It says nothing about a purchase that happens anywhere else.

The fastest way to judge the method is to see it run.

The example dashboard is a fictional brand, labelled as one throughout, so you can read every surface without a run of your own.

Book a demoWhy we exist