Skip to content

AnswersWrong product answers

When an assistant tells a buyer your product is discontinued, out of stock, or the wrong price

Four commerce errors end a sale before anybody reaches your site: a live product called discontinued, a price that is not yours, an availability claim your own pages contradict, and a buyer sent somewhere they cannot buy. Each one comes from a different source, and only some of those sources are yours to correct. The first job is to work out which.

What kinds of wrong answer actually cost you the sale?

Four commerce errors end a purchase: an assistant calls a live product discontinued, quotes a price that is not yours, states an availability your own pages contradict, or sends the buyer somewhere they cannot buy it. Each is reported separately because each is fixed by a different person.

These four sit in the last of the five gates, the one that asks whether a buyer sent by an assistant can complete the purchase at all. They are kept apart from each other because a wrong price and a wrong shop are different failures with different owners, and rolling them into one number tells nobody what to do on Monday.

  • Discontinued. The most expensive of the four. A buyer told a product no longer exists does not go looking for it, does not open your site, and does not ask anybody. There is no second chance inside that answer.
  • Wrong price. Sometimes a real error and often not, which is why it has the most careful rule of the four. More on that below.
  • Wrong availability. Out of stock, hard to find, or not sold in this country. A buyer reads it as a reason to look at the next name on the list.
  • Wrong place to buy. The answer recommends you and then names a shop that does not carry you, or names no shop at all. A named retailer is worth more to a buyer than a marketplace, and a marketplace more than being sent to your own site.

Where does a wrong price or a discontinued claim come from?

A wrong commerce answer comes from your own pages, from a retailer or marketplace listing, from a third-party page such as a review or a forum thread, or from what the model already held about your category. Which of those it is decides whether you can correct it, influence it, or only earn your way into it.

Sorting those sources by how much of them is yours is the whole of the diagnosis, and it is the step most teams skip.

  • What you control. Your own product pages. Where the price, the availability and the product identity are missing or stale in the markup, an assistant fills the gap from somewhere else.
  • What you can directly influence. A retailer listing. You do not own the page and you have a trading relationship and a named person who can change it.
  • What you can attempt to earn influence over. An independent publication, a review site, a comparison page. You can supply better facts. You cannot demand a correction.
  • What you cannot control. An organic forum thread, and whatever the model already held about your category before your last change. Neither is a page you can edit.

The order matters twice. The cheapest correction is almost always the innermost one, and the outermost one is where any product promising to fix everything stops being honest with you.

A colleague sent me a screenshot. Is that enough to act on?

A screenshot is one answer, from one assistant, on one day. These models do not answer identically twice, so a single wrong answer tells you the error is possible rather than how often a buyer meets it, and how often is the number that decides what the fix is worth.

The same question put to the same assistant twice can produce two different answers, and put to five assistants it routinely produces a disagreement. So the reading worth acting on is a rate rather than an instance: how many of your buying questions produce the wrong claim, on how many assistants, and whether it is still there next month.

Delphi puts your category's real buying questions to ChatGPT, Gemini, Claude, Perplexity and Grok, and asks each of them more than once. How far the five disagree with each other on the same question becomes a confidence band that travels with every figure, and a movement smaller than the band is stated as a movement nobody can stand behind yet.

That cuts both ways, and the second half is the useful one. A wrong claim that appears once and cannot be reproduced is not a crisis. One that appears on four assistants and is still there on the next measurement is.

How do you find every wrong answer instead of the one somebody noticed?

Delphi puts a brand's real buying questions to five assistants, stores every answer word for word, and compares the commerce claims inside those answers against what the brand's own product pages publish. A claim can then be quoted back with the assistant that made it named beside it.

Storing the answer verbatim is what makes a finding usable outside your own team. A retailer, a publisher or a colleague who never saw the screenshot can be shown the exact sentence, the assistant that produced it, the model that produced it, and the sentence on your own page that says otherwise.

Where the two disagree, both sides are quoted and neither is declared correct. That is deliberate rather than timid. A buyer asking an assistant never opens your page, so it makes no difference to the lost sale whether your page is right and the answer is stale, or your page is out of date and the answer is reading it correctly. Either way there is work to do, and presuming your page is the truth would send half of those cases to fix the wrong end.

Where your page says nothing on the subject, that is recorded as no comparison rather than as agreement. Accuracy is only scored where something was genuinely comparable, so a claim nobody could check never counts in your favour and never counts against you.

What can you actually fix, and what is out of reach?

Your own pages are fixable this week: publish price, availability and a stable product identity in machine-readable form, and keep them current. A retailer listing is fixable through whoever owns that account. A third-party page can be earned into rather than instructed, and what a model already holds about your category cannot be edited at all.

The limits are worth stating before the capabilities, because a plan built without them sends people to work on things that will not move.

A model can change underneath you. An update can move what an assistant says about you with nothing changing on your side. Every stored answer records the exact model that produced it, so a movement caused by a model change is visible as one rather than mistaken for your own progress.

An answer shaped by one person's own history is not reproducible, and Delphi does not try to model it. What is measured is the answer a new buyer with no history receives.

The pages read are your own, not a retailer's shelf. Where an assistant says a shop stocks you, that is the assistant's claim rather than a checked listing, and it is reported as a claim. Stock and shelf position inside a specific retailer are not measured here.

The assistants retailers run inside their own shops are not read. Nothing here measures those, and nothing here implies otherwise.

The assistant quoted a lower price than we publish. Is that an error?

A quoted price counts as wrong only where it diverges materially from what you publish, and the band is 25 percent either way. Retailers discount, run promotions and sell different pack sizes, so an assistant quoting a real shelf price below your list price is describing the market accurately rather than making a mistake.

A tight band manufactures problems. Set the tolerance at a few percent and every promotion in your category becomes a defect on somebody's list, the list stops being read, and the one answer quoting half your price is buried inside it. What a generous band catches is the divergence that changes a buyer's decision.

Two smaller rules sit underneath it, and both exist to stop an invented comparison.

  • A quoted price is compared against the closest price you publish, because a brand sells more than one thing and an answer about the small pack must not be judged against the price of the large one.
  • A price with no digits in it is not read as a number at all. An answer saying around fifteen pounds is not turned into fifteen, because comparing against a figure nobody stated is exactly the fabrication this measurement exists to catch in others.

Where your pages publish no price, price accuracy is not scored as zero. It is reported as unmeasured and its weight moves to the parts that could be read, because a brand nobody can check is not a brand that got it wrong.

What should you do first?

Establish how often the wrong claim appears before changing anything, then work outward from what you control: your own pages first, the retailer listings you can reach second, and the third-party pages the assistants are visibly reading third.

The order is not arbitrary. Your own pages are the source an assistant can read most cheaply, they are the fastest thing you can change, and the same markup that settles a price claim usually settles the availability claim sitting beside it.

Then read the answer for what it named. An assistant that sends buyers to a particular shop is telling you where the correction has to land, and the person who can make it is usually whoever owns that account rather than anybody in marketing.

Everything past that is earned rather than instructed. A comparison page that has your product wrong is worth writing to with the correct facts, and the answer may be no. That is a real limit and not a reason to skip the letter.

Then measure again. A claim that appeared on four assistants and now appears on one is a result you can take to a board. A claim that moved by less than the measurement's own margin has not yet moved.

Which of the five is holding your brand back?

A wrong answer about your own product is cheap to find and slow to unlearn, so the useful first step is knowing how often a buyer actually meets it.