Skip to content

AnswersContradicted claims

When an assistant and your own page tell a buyer different things

A contradiction is a comparison between two published things: what an assistant told a buyer, and what your own page says on the same subject. Delphi reports the disagreement, quotes both sides, and rules on neither. Most claims cannot be compared at all, and those are reported as unmeasured rather than scored either way.

What makes a claim comparable to your own page?

A comparable claim is one your own published page addresses specifically, so the two either agree or say something different: a price, a stock state, a specification, a product that exists or does not. A claim your page does not address, or addresses only vaguely, is not comparable and is never recorded as a disagreement.

An assistant makes dozens of assertions about a brand across a set of buying questions. Only some of them can be set against anything. The test is narrow on purpose: your page has to address the same specific thing, and then either agree with the assistant or differ from it.

  • Comparable. An assistant quotes a price and your page publishes one. It says a product is no longer made and your page lists it as available. It states a size, material or compatibility your page states differently. It describes a product or range that appears nowhere on your site at all.
  • Not comparable. An assistant says a rival suits sensitive skin better. It says you are the premium option. It describes a shop's shelf rather than your own. It asserts anything your site is simply silent about.

The second list is the larger one for almost every brand, and that is the ordinary case rather than a failure of the run. Silence is never read as a disagreement, and an assertion that merely seems implausible is not one either. Only a specific difference against something you published counts.

Why is a claim your page never mentions reported as unmeasured?

A claim nothing could be checked against is reported as unmeasured because scoring it either way would state a finding nobody established. Counted as agreement it flatters the brand, and counted as an error it blames the brand for a page that was simply silent. Factual accuracy is only scored where at least one claim could be compared.

This is the constraint that decides how much of the rest is worth anything, and it is the part most measurement omits. A score that starts perfect and only ever falls when something is caught reports a perfect score for an envelope nobody opened. Every brand an assistant has heard of reads as accurately described, forever, and nobody can tell that from the real thing.

So where no claim in a run could be compared, factual accuracy is not scored at all, and the gate it belongs to rests on how the brand is described instead. Where nothing at all could be read from the brand's own site, every claim in the run comes back as no comparison, and nothing about accuracy is claimed in either direction.

The same rule runs through the model above it. A gate that could not be measured is excluded and the remaining weights are renormalised, and how many gates were scored is published beside the figure. Nothing is ever counted as zero for not having been looked at, because a zero says the brand failed at something and caps a score it cannot move.

Who decides whether the assistant or your page is right?

Nobody rules on it. Both sides are quoted, the assistant's sentence and your page's own wording, and neither is declared correct, because a buyer reading the assistant never opens your page, so the sale is lost the same way whether your page is right and the answer is stale, or your page is out of date and the answer is reading it correctly.

This reads at first like a product refusing to do the hard part. It is the opposite. Adjudicating truth needs standing that nobody measuring answers actually has: it would mean deciding, from outside your company, whether a specification changed last quarter or a line was retired, and being wrong about that in front of your team.

It is also the wrong question commercially. Presuming your page is the truth sends half of these cases to fix the wrong end. If an assistant tells a buyer a product is discontinued, that purchase does not happen. Whether the correct move is to reach the sources the assistant is reading, or to correct your own page, is decided by evidence about where the claim came from and not by a verdict about who was right.

So a finding is written as one plain sentence carrying both halves: the assistant says this, your page says that. Your page's own wording is quoted beside it as one side of the disagreement rather than as a ruling, and the assistant's sentence is stored word for word so it can be shown to a retailer, a publisher or a colleague who never saw it.

How is a contradiction found across five assistants rather than in a screenshot?

Delphi puts a brand's real buying questions to ChatGPT, Gemini, Claude, Perplexity and Grok, asks each of them more than once, stores every answer word for word, and pulls the claims out of those answers. Identical claims are grouped across assistants and repeats and checked once, and the finding names which assistants made it.

These models do not answer identically twice. A screenshot tells you a wrong answer is possible; it cannot tell you how often a buyer meets it, and how often is the figure that decides what the fix is worth. One assistant saying something once and four assistants saying it every time are different problems with different budgets.

So the same assertion arriving from several assistants and several repeats is treated as one claim and checked once, against the same evidence, and the check is then applied to every answer that made it. That is cheaper and it is also the right shape for a reader: a claim four assistants agree on is one thing to fix, not four.

The finding then names the assistants rather than counting them. Two of five is not a next step. Perplexity and Gemini carrying a claim that ChatGPT does not is a signal about what those two are reading, and that is usually what decides whether the work is on your page or somewhere else entirely.

The checking is budgeted, and the budget goes to the claims most assistants repeat first. A run that produces more distinct claims than the budget covers leaves a tail of the rarest ones unchecked. An unchecked claim never counts as agreement and never enters the accuracy figure, so the basis is smaller than the claim count rather than quietly wider.

Five findings, all answered by one sentence on our site. Is that five problems?

One problem with one fix. Findings are grouped by the wording on your own site that answers them, so a page already saying the right thing and five answers saying the opposite is a single problem, and the repetition is the finding rather than a list of five.

This was learned by getting it wrong. Drawn one claim at a time, the same sentence from a brand's own site printed five times running, because one page answers five separate assertions. It reads as a rendering fault and it makes five entries of somebody's work.

Grouped by the wording that rebuts them, the repetition becomes the point. The number of assistants repeating a claim your page already answers is a measure of how far your own correct sentence is failing to travel, which is a much sharper problem than five unrelated ones and a much cheaper one to act on.

It also sorts the work before anybody reads a card. A group whose rebuttal is already published is the cheapest thing on the page, because the correct wording exists and the job is getting it read. A claim your site answers nowhere is a different job, and it is not made cheaper by sitting in the same list.

What if the assistant is right and it simply looks bad?

An unflattering claim your own page does not contradict is not a contradiction at all. It is reported as no comparison and never counted as an accuracy error, and the work it points at is evidence or the product itself rather than a correction anybody can ask for.

The two cases need separating early, because they take different people different amounts of time. A false claim gives you something to point at: a published sentence of your own, and a set of sources the answers were visibly reading. A true one gives you nothing to point at, and asking a publisher to remove an accurate sentence is not a plan worth writing down.

An unflattering claim still counts, in a different place. How favourably, neutrally or unfavourably the answers speak about a brand is measured separately from whether they are accurate, so being described badly and being described wrongly are two readings rather than one. A named and criticised mention is treated as worse than a neutral one and better than not being mentioned at all.

The honest edge, and it runs the other way. The comparison records a difference whichever side is the optimistic one, so a brand whose own marketing claims more than its evidence supports collects findings against it. There is no separate verdict for a claim your own page makes that the assistants dispute. That is a known gap in the model rather than a judgement about your copy, and it is stated here because a finding read as an assistant error when it is really your page being ambitious would send somebody to fix the wrong thing.

What do you actually do about one?

Establish how many assistants carry the claim and which sources those answers cite, then work outward from what you control: your own pages first, the retailer and publisher listings you can reach second, and measure again, because a change smaller than the measurement's own margin has not yet moved.

Your own pages come first because they are the source an assistant reads most cheaply and the only one you can change this week. A price, a stock state and a stable product identity published in machine-readable form settle several claims at once, and their absence is what an assistant fills from somewhere else.

Then read what the answers cited. A retailer listing is somebody else's page and you have a trading relationship and a named person who can change it. An independent publication or a comparison page can be written to with better facts, and the answer may be no. An organic forum thread is not a page anybody can edit, and neither is whatever a model already held about your category. Sorting a finding into those four is most of the decision about who does the work.

Then measure again, and be strict about what counts as a result. The measurement carries a margin drawn from the disagreement between its own repeated answers, and movement smaller than that margin is reported as movement nobody can stand behind. A claim carried by four assistants that is now carried by one is a result. A claim that shifted inside the noise is not, and calling it one is how a team spends a second quarter on a fix that never worked.

Which of the five is holding your brand back?

A contradiction is cheap to find and slow to unlearn, so the first useful step is knowing how many buyers meet it and which assistants are carrying it.