AnswersTracking over time
How to track brand mentions in AI assistants, and tell a real change from noise
Tracking what ChatGPT and the other assistants say about your brand means asking the same prompts the same way, several times each, on a regular rhythm, and recording the same things every time. Repetition is what makes it tracking: in Delphi's own study of answers collected up to 24 August 2026, an assistant asked one question more than once inside a single measurement named the same set of companies only 21.5% of the time. One look shows what an assistant can say. Repeated looks show what it usually says, and whether that is moving.
Why is one check not enough to monitor what ChatGPT says about your brand?
One check is not enough because the same prompt usually does not return the same set of names twice: in Delphi's study of answers collected up to 24 August 2026, an assistant asked one question more than once inside a single measurement named an identical set of companies only 21.5% of the time.
Why the answers move, and what one look can still tell you, is covered in how to find out what ChatGPT says about you. For tracking, it means each point on the line has to be a rate. A series you can keep by hand looks like this: a fixed list of buyer prompts, each asked several times of each assistant you follow, every answer logged as its own row. The figure for the period is the share of rows that named you and, separately, the share that recommended you, which are two parts of what AI visibility covers.
Keep each assistant as its own series. Of the companies named anywhere in that study, 75.7% were named by exactly one of the assistants, so what one assistant says is not a reliable guide to what the others say.
What should you record so two checks can be compared?
On top of whether you were named, whether you were recommended and who else appeared, record where you sat in the list and the exact prompt, the date and the model that answered. Position shows a slide that a count of mentions misses, and the other three let you rule out a change in how you asked before you read a change in the market.
The fields for a single look are on how to find out what ChatGPT says about you. Three additions matter once looks are compared:
- Position. The recommendation, one of the first three names, or further down. A slide from first to third leaves a count of mentions exactly where it was.
- Steered away from. An answer that names you and then advises the buyer against you still counts as a mention in a simple tally. Keep it apart from the first day, for the reasons in mentioned is not recommended.
- The prompt, the date and the model. The exact wording, the day, and the model name where the product shows one. Without them, a change in how you asked and a change in the market look the same on the line.
Delphi keeps every answer word for word, with the exact prompt it sent and the model that replied. For each answer it reads, it records whether the brand was the top recommendation, named, named but advised against, or absent, its position, how it was described, the rivals named with their own positions, and the sources cited where the assistant returned them. The score is built only from the answers it has read.
How often should you track brand mentions in AI search?
How often matters less than how many askings sit behind each figure you compare. A monthly reading built from many askings and a month of daily askings pooled into one rate are both usable, while a few prompts read one day at a time mostly record the assistants' own variation.
The repeat figure above is the reason. A difference that shows up between two askings in one sitting cannot be read as a movement when it shows up between Monday and Tuesday, so a few prompts, each asked once and read as that day's result, will show changes on many days whether or not anything has moved.
The real choice is how to spend the askings. One large reading a month, with every prompt asked several times of each assistant, costs one sitting and carries its own spread, but it cannot say when within the month something moved. The same askings spread across the month and pooled into one monthly rate are just as valid, and can be split by week to show roughly when a change began, with a wider spread on each week. Either way, keep the prompts fixed, and take an extra reading after anything that could move an answer, such as a launch, a new comparison page or a model release.
How do you know a change in AI answers is real?
A change in what AI assistants say about your brand is real only when it is larger than the spread their answers show with nothing changed. Build each reading from several askings of every prompt, take the baseline more than once to see how far readings sit apart on their own, and treat a later shift inside that range as no confirmed change, however good or bad it looks.
Measure the noise before you measure the change. Take your baseline more than once, a few days apart, with nothing changed on your side: the same prompts, each asked several times. The spread across those readings is a first estimate of the noise, and it firms up as readings accumulate. A move from one month to the next that stays inside it is not a result, and reporting it as one sends a team to explain a change that did not happen.
Compare like with like: the same prompts, asked the same way, of the same assistants. A rate built from many askings and a rate built from a few are not equally reliable, even when they print the same percentage. The same holds when checking whether work on ChatGPT landed: the reading before the work and the reading after it have to be taken the same way.
Delphi uses a different spread for the same purpose: the band published with its score is how far the assistants disagreed with each other in that run, and its dashboard and history page report a movement smaller than the band as no confirmed change. How that band is computed is on how many questions a measurement needs.
Did your AI visibility change, or did the way it was measured?
To tell a change in AI visibility from a change in how it was measured, check the method before the market: if the prompts, the session they were asked from or the scoring rule changed between two readings, the step between them describes how you measured rather than what buyers are told. If the model behind the assistant changed, the step may not be yours, because a model update can move answers with nothing changing in your market.
- The prompts. A reworded prompt is a new question. Keep the old wording, or start a new series and do not join it to the old one.
- The session and the place. Ask from a clean session, from the same place, every time. OpenAI documents that, where Reference chat history is switched on, ChatGPT can use information from past conversations to personalise later responses, that a temporary chat set to Unpersonalized does not use existing memories, and that ChatGPT may use an approximate location from your IP address for local results. In ChatGPT, a tracker run from one signed-in account, or from wherever somebody happens to be that day, may be partly measuring that account and that place.
- The scoring rule. If what counts as a recommendation changes halfway through, the line changes with it. Write the rule down on the first day.
- The model. OpenAI's release notes record changes to the models in the ChatGPT app, including updates to a model that keeps its name and the retirement of older ones, and say which of them also apply to the API. A movement that lines up with one may come from the assistant rather than from last month's work, so check which model answered before crediting or blaming the work.
Delphi holds these still or records them. Each answer is requested as a single new question with no conversation before it. Editing a question set creates a new version rather than rewriting the old one. Every answer is stored with the exact model that produced it, and the download for each run carries that model beside every answer, so for a Delphi reading that stored model is the thing to check rather than an app's release notes. And every run records which version of Delphi's own reading rules labelled it: where the questions or those rules changed between two runs, the trend view says so above the chart, so a step made by the method can be told apart from one made by the market.
Do you need a daily AI brand monitoring tool?
A daily tool is worth having only if its daily figures rest on enough askings to carry their own spread, or are pooled into weekly or monthly rates before anyone reads them, because a few prompts asked once a day mostly record the assistants' own variation. Delphi is not a daily tracker: it measures in runs, once or twice a month for each brand depending on the level, and each run is set up to put every question in the set to ChatGPT, Gemini, Claude, Perplexity and Grok more than once.
The runs each level includes are on the pricing page. Once a brand has its first measurement, it is measured again automatically about once a month, when its last run is at least four weeks old and the month's allowance has a run left, unless automatic runs are switched off for that brand. The automatic run is one a month; where a level includes more, the others are started by hand. A month that needs one more reading, after a launch for example, can have one for $49.
What goes into each run is the point. A question set holds up to 40 questions, every question is planned for every assistant more than once, and how far the assistants disagree becomes the band published with the score. A run that reaches fewer assistants than planned says so rather than filling the gap.
Related answers
- How AI visibility is measuredHow the score is built, including what it cannot see.
- How to find out what ChatGPT says about youThe free check by hand, which comes before any tracking.
- How many questions a measurement needsHow large each reading has to be before a change in it means anything.
- What AI visibility tracking costsWhat repetition and coverage cost at each size, starting with the free version.
Which of the five is holding your brand back?
Much of what changes between two quick checks is the assistant answering differently, so each point on a tracking line needs enough askings behind it to be told apart from the next.