Short answer

A generative AI does not return a stored result the way a classic search engine does; it generates an answer each time, with some randomness, with whatever index it has that day, with the context of your session and, on some platforms, with your location. That is why two people see different lists, and why you see a different list on Tuesday. This does not make measurement useless: it makes repetition compulsory. You measure with the same questions on different dates and label each combination as stable, variable or not assessable.

Five reasons it changes

  1. Model randomness. Models generate text by sampling among likely options. Two identical runs can produce two wordings and, sometimes, two lists.
  2. The search index. When the platform searches the web, it retrieves what is there that day. A new page, an updated listing or a news story changes what gets retrieved.
  3. Session and memory. ChatGPT can use previous conversations and memory to personalise. If you have asked about your company before, you are no longer measuring your market: you are measuring yourself.
  4. Location and language. Asking in Spanish from Madrid and in English from Dublin are two different queries. Google AI Overviews is especially sensitive to this.
  5. Wording. “Best online accountant” and “which online accountant would you recommend?” can give different lists. That is why the exact text of each question is frozen and versioned.

What it means for you

  • A single screenshot describes a moment. It cannot tell you “we do not appear” or “we appear now”.
  • A difference between two screenshots does not prove a real change until it repeats.
  • An aggregate rate (“we appear in 40%”) hides which rows changed. The full matrix matters more than the summary.

How to measure anyway

Freeze the questions. Exact text, version, date. If you change a word, it is another question.

Use a clean chat. A temporary chat or a session with no history or memory. Note whether web search was on.

Fix the platform, language and location. And record them in every observation.

Repeat with a gap. Two rounds seven days apart are the minimum to talk about stability. Three or four rounds start to separate trend from noise.

Label each combination of question and platform:

LabelMeaning
StableSame mention (or same absence) and similar main sources in every round
VariableThe mention, the order or the main source changes between rounds
Not assessableAn error, a block or an insufficient answer in some round

Record errors as errors. A timeout is not a zero. It is a failed observation that you run again.

How to read a variable result

If your brand is variable in half the questions, there are two possible readings: the category is not yet “decided” in the AI, or your presence depends on sources that are sometimes retrieved and sometimes not. In both cases the action is the same: strengthen what you control (pages that answer, a consistent entity) and keep measuring before investing in external sources.

What variability does not justify

It does not justify giving up on measuring (“it changes all the time anyway”), nor hiring whoever promises you stability. It justifies exactly the opposite: measuring with a method and distrusting anyone who shows you a single screenshot.