Skip to content

AnswersWhere to start

Which part of the journey is failing, and what to fix first

The order comes out of the measurement rather than out of a template. Five things have to happen before a buyer buys, they happen in sequence, and the weakest one caps every step after it, so that is where the work starts. Fixing a step you already pass changes nothing downstream, however much it improves.

What has to happen before a buyer buys?

Five things have to happen in order before an AI assistant sends a buyer to a brand: the assistant has to be able to read the brand, name it in answers, describe it accurately, put it at or near the top, and tell the buyer where to get it.

This is not a marketing funnel. It is the order an assistant works in when somebody asks it what to buy, and each thing depends on the one before it.

  • Found. Can the systems behind the assistants reach and read you at all.
  • Seen. Do you turn up when a buyer asks a question about your category.
  • Trusted. Is what the answer says about you accurate, and does it read favourably.
  • Chosen. Are you the recommendation, or one of several names in it.
  • Bought. Does the answer end somewhere a buyer can act on.

Two things follow from the order. A brand can be excellent at four of the five and sell nothing. And a brand can move one step and watch the ones after it move without being touched.

Why does fixing your strongest area first waste the money?

Fixing the strongest step first wastes the money because the five steps run in sequence: work on how a brand is described earns nothing while no answer is naming it, and a buyer who is never shown the name never reaches the question of whether to trust it.

Whatever a later step scores on its own, the part of it above the weakest step is unreachable until the weakest one moves. So the practical worth of a gain on your strongest step is close to nothing, and the same gain on the weakest one is worth itself plus whatever it releases behind it.

That is also why an ordered plan is not the same as a ranked list of problems. Ranking by how large each weakness looks puts the biggest gap first, and the biggest gap is very often on a step that nothing downstream is waiting for. The measurement knows which one is holding the rest, and a template cannot.

Delphi names the weakest step the binding constraint, for the plain reason that it binds the four after it.

We never appear at all. Where does that come from?

Never appearing is usually an access problem before it is a content problem: if the crawler behind an assistant is refused, or the facts a buyer needs only exist after JavaScript runs, there is nothing for that assistant to read and every step after it is capped for that assistant.

Four things decide whether a brand can be read. Whether the named crawlers behind the assistants are allowed. Whether pages carry the machine readable labels that say which text is the price, the rating, the ingredient list. Whether the facts a buyer needs are in the HTML as delivered, rather than appearing after a tab is clicked. And whether the assistants already hold the brand as a distinct thing rather than an unfamiliar string of letters.

It is the cheapest of the five to fix and the easiest to miss, because a refused crawler is invisible from the inside. Your analytics show the site working. The readers it is not working for never arrive, so they are not in the report.

Delphi identifies itself by name to every site it reads, and reports a block as a finding rather than working around it. A measurement that routes around a refusal would be measuring a site the assistants cannot see.

Being named and not recommended is a different failure from being invisible, and it is measured as two separate things: appearing in an answer at all, and being the recommendation rather than one of several names inside it.

The distance between those two is demand you have already paid to create, arriving at the answer and going to a rival. Every pound spent on awareness has worked, and the last step of it has been handed to somebody else.

Two readings sit underneath. How often you are the first name, and how often you reach the first three, which is the set a buyer actually considers. Both count only questions asked at the point of buying. An informational question has no recommendation to win, so it stays in the set and is excluded from this score rather than manufacturing a zero nobody could act on.

A brand named first and then warned against is not counted as a recommendation. An answer that opens with the one most people buy and then says it would not is a Trusted problem wearing a Chosen result, and counting position alone would score the most damaging answer in the run as a win.

The work here is usually not on your own pages. An answer is assembled from what the assistant read, so the useful output names the third party pages that were cited on the buying questions you lost, and how many of those questions each one touched.

An assistant says something untrue about us. Is that the thing to fix?

An untrue claim is a Trusted problem, and Delphi quotes the assistant's wording beside the brand's own published wording without ruling on which is right, because the buyer asking the assistant never sees the brand's page.

That refusal is deliberate. If an assistant tells a buyer a product is discontinued, the sale does not happen, and it makes no difference to that outcome whether your page is right and the assistant is stale, or your page is out of date and the assistant is reading it correctly. Either way you are losing the sale and either way there is work to do. Presuming your page is the truth would send half of these cases to fix the wrong end.

Where your own pages do not address a claim at all, that is recorded as nothing to compare rather than as agreement, and accuracy is scored only where something was genuinely comparable. A brand nobody could check is never reported as a brand that got it right.

Severity is weighted rather than counted, because one invented negative can undo a whole content campaign and a wrong pack size cannot. Whether this is the thing to fix first still depends on the sequence: an accurate description earns nothing while no answer is naming you.

Buyers hear about us and cannot get anywhere to buy. What is failing?

The last step is Bought, and it reads the end of the answer rather than the brand's website: whether the assistant names somewhere to get the product, how specific the offer is, and whether it raises something that stops the purchase.

Auditing a brand's own storefront is the obvious way to measure this, and for most brands it is wrong, because most brands do not sell on their own site. It scores a shop that does not exist, and a number that reads near zero for everybody has stopped discriminating between anything.

What the assistant says at the end of an answer is the transaction path as the buyer experiences it. Where it sends them. Whether the offer carries a price, a pack size or a named variant rather than the brand name alone. Whether it tells them the product is discontinued, out of stock, hard to find or not sold where they are. And, weighted lowest of the four, whether your own site is ready for an agent to transact against.

Where a comparison cannot be made at all, for instance where your pages publish no price, it is reported as unmeasured and its weight moves to the parts that could be read. Not knowing outscores knowing it is wrong, which is the right way round.

Why one place to start rather than a list of weaknesses?

One place to start is more useful than a list of five weaknesses because a list is not a sequence, and a team handed five weaknesses starts with the one it already knows how to fix rather than the one holding the rest back.

The weakest step sets the order of the plan. It does not decide the whole plan, and it should not. Whether a brand can move a given step depends on circumstances outside its control, so a product that points at the weakest number and stops has told a commercial director to go and do something they may not be able to do.

So the rest is presented as a ranked set rather than a queue, and it is prioritised on four things that are all real stored figures: what the work would move, how much room is left on the step it targets, how much confidence the evidence behind it supports, and how much effort it takes. An item weak on any one of those is not carried by being strong on another.

The sequence is a claim about how the five steps interact, and it is checkable rather than asserted. If the first piece of work ships and the step it targeted does not move by more than the measurement's own margin of error, that is reported as no confirmed change rather than as a win.

What does a plan derived from a measurement look like?

A derived recommendation carries a trigger that was evaluated against a signal the run actually measured, a named target, the mechanism it works by, an effort estimate, and an expected lift stated as a range with the basis it rests on.

The trigger is what makes it derived rather than templated. It is evaluated in code rather than asked of a model, and there is no close enough: a signal the run did not measure resolves to false, so a thin measurement produces a short plan instead of a confident wrong one.

  • Trigger. The condition that fired, and the measured signal it was evaluated against.
  • Target. The page, the domain or the part of your own site the work acts on, named rather than described.
  • Mechanism. Why it works, and which of the five steps it is trying to move.
  • Effort. A stated band with its reason attached. Nothing here measures effort, so the reason travels with the band rather than a bare letter standing in for one.
  • Expected lift. A range with its basis, never a point estimate.

Every recommendation also states what would count as it having worked: what a next measurement would have to see, and what a next measurement still could not see whatever happens. Work ordered this way respects its own prerequisites, so nothing is scheduled before the thing it depends on.

The same inputs give the same plan. There is no sampling and no clock in the part that chooses, which is what makes a comparison between one month and the next mean anything at all.

Which of the five is holding your brand back?

A sequence is only worth as much as the measurement under it, which is why every answer these steps rest on is stored word for word and can be read back.