In this article 13
- Start with your own free data
- 1. Start with a decision
- 2. Design the question set
- 3. Set the scope
- 4. Keep the evidence
- 5. Classify without hiding the sample
- 6. Show the variability
- 7. Separate observation, inference and action
- 8. Prioritise by control and evidence
- 9. What the client should receive
- Common mistakes
- Minimum template
- When to turn it into tracking
Short answer
To measure a brand’s visibility in AI answers (ChatGPT, Gemini, Claude, Google’s AI Overviews) you need a versioned set of questions, more than one platform, more than one round and the original answer from every run. Then you separate presence, recommendation, accuracy, sources and variability. You do not need to start with a score. The smallest unit is not a position; it is a traceable observation: question, platform, date, region, language, full answer and visible sources.
Start with your own free data
Before you pay for a dashboard, check what the brand’s verified properties already show. Google Search Console offers dedicated reports for the generative AI features in Search, such as AI Overviews and AI Mode, with impressions by page, country, device and date. Bing Webmaster Tools offers, in public preview, citations, cited pages, trends and grounding queries (the key phrases its AI used to retrieve the content it cited) for Copilot, AI summaries in Bing and partner integrations.
They are valuable sources, and nobody should hide them to justify a service. They also have limits: Google does not document the literal question or the full answer in that report, and Bing states that its data is a sample of overall citation activity. Neither replaces a measurement with exact questions, saved answers, competitors and two comparable rounds.
1. Start with a decision
A measurement with no decision behind it produces a nice dashboard and little useful work. Define a single business question:
- does the brand make it into comparisons when the user does not name it?
- is it described correctly?
- which sources appear for the category?
- which competitor dominates a specific intent?
- which gap can your own website close?
- is ongoing tracking worth it?
The decision sets the questions, the market, the platforms, the rounds and the deliverable.
2. Design the question set
Each question should record a stable ID, the exact wording, its origin, the intent, the buying stage, the language and country, whether it includes the brand name, the version and the date it was added.
Origin
Prioritise questions that come from real interviews and support tickets, internal site searches, sales conversations and query research. Bing’s grounding queries are a clue, not a literal question. An explicit hypothesis is acceptable when there is no other source, as long as it is not presented as observed demand.
Branded and unbranded
“What do you know about Nubaria?” is a branded question: it measures recognition or representation. “Best invoicing software for freelancers” is an unbranded question: it measures discovery. Mixing them into a single rate inflates the reading.
3. Set the scope
Our AI diagnosis uses 12 questions, 2 platforms and 1 round: 24 observations. The ongoing partnership repeats 20 questions on 2 platforms twice a month. The number does not prove representativeness; it makes the work, the cost and the gaps visible.
Before you start, record: platforms and visible versions, web access on or off, clean or signed-in session, region and language, the interval between rounds, the retry limit, how errors are handled and which data must never be entered into an assistant.
4. Keep the evidence
For each observation, save: run ID, question and version, platform, date and time, status (valid, error or unavailable), full answer, brands present, explicit recommendation if there is one, position if there is a list, verifiable claims, visible sources, extraction method, confidence and reviewer.
A screenshot is useful for checking layout and status. It does not replace a structured record or a definition.
5. Classify without hiding the sample
Mention rate
Valid answers that mention the brand, divided by eligible valid answers. Publish the numerator and the denominator. Exclude errors by a rule defined before you see the result.
Recommendation rate
Valid answers that explicitly recommend the brand, divided by eligible valid answers. A mention is not a recommendation: “we do not recommend Nubaria” contains the brand and is not a favourable result.
Position
It only exists when the answer presents an order you can interpret. Do not give a position to a paragraph with no list, and do not turn the first name in a sentence into a ranking.
Accuracy
Check specific claims against a source: correct, partial, incorrect or unverifiable. A positive tone does not make a claim correct.
Sources
Count domains and URLs separately. A visible source does not prove causation. Record its type: your own, a competitor’s, a third party’s, institutional, documentation or community.
6. Show the variability
For each question and platform pair, compare rounds: present in all, absent in all, variable or not assessable. Compare facts, recommendation, position and sources too. Two answers can both mention the brand and describe it differently.
An aggregate rate can stay stable while almost every row changes. That is why the full matrix must be available, not just the summary.
7. Separate observation, inference and action
| Type | Content |
|---|---|
| Observation | A comparison site’s guide appeared in 9 of 24 answers |
| Inference | It may be a relevant source for that intent |
| Action | Check whether your listing there is correct and provide a legitimate update |
| Verification | Repeat the same questions after the update |
Do not write “the guide caused the answers” if the protocol cannot prove it.
8. Prioritise by control and evidence
Each action should carry the finding behind it, the expected mechanism, the client’s control from 1 to 5, the strength of the evidence from 1 to 5, the cost, the owner, the date, the risk and the observation that will assess it later.
Work with high control and strong evidence goes first. An external, expensive action based on a single answer waits.
9. What the client should receive
Executive summary, scope and methodology, presence by intent and platform, claims and accuracy, source map, variability, action plan, full matrix, error log, data export, definitions and version, limits and non-guarantees.
A score without an appendix cannot be audited. An appendix without a summary cannot be acted on. You need both levels.
Common mistakes
- Asking once. The answer can change. One round does not let you talk about stability.
- Writing questions that produce the result you want. Branded questions or leading statements inflate presence.
- Mixing platforms. An average can hide that the brand appears on one and not on the other. Show both.
- Treating errors as zeros. A timeout does not prove absence. Log it as an error.
- Calling a sample market share. A selection of questions is not a representative survey. Say “observed rate in the sample”.
- Attributing the change to the latest action. A different answer can coincide with a new page without being caused by it.
- Hiding the original answer. Without evidence, the client can only trust the provider. That dependence lowers the value of the report.
Several of them are covered in more depth in mistakes when measuring AI visibility.
Minimum template
Decision · market and language · questions and their origin · platforms · rounds · planned observations · errors · mentions · recommendations · incorrect claims · main sources · variable combinations · three actions under your control · next round.
Copy this list into a spreadsheet and you have the skeleton of a baseline. The rest is discipline: same questions, same platform, a different date.
When to turn it into tracking
Do not pay for recurring tracking until you have shown that someone needs a record over time. It makes sense when there is a baseline, there is a recurring decision, the question set stays stable, the company will act on changes and a second round has already produced value.