In this article 5
Short answer
An AI answer can come from two places. From what the model learned during training, which cites nothing because it does not “remember” where it learned it. Or from a live search, in which the platform retrieves pages from the web, reads them and generates an answer citing some of them. Only the second route shows sources, and it is the one you can work on: the pages it cites are crawlable, clear and relevant to the question at the moment it is asked.
Two mechanisms, two behaviours
Without search. The model answers with what it learned. If your company was not relevant on the web when it was trained, you do not exist for it. There are no sources and no way to “update” it until a new version comes out.
With search. ChatGPT with search, Perplexity, Gemini with web access, Google AI Overviews and AI Mode, Copilot. The platform queries an index, retrieves documents and the model writes with them. Here there are sources, and they change when the pages change.
When you measure, always note whether search was on. Comparing an answer with search and one without is comparing two different products.
Who crawls what
Each platform has crawlers for different purposes. The search crawlers are the ones that matter for appearing:
| Platform | Search crawler | Training crawler | Note |
|---|---|---|---|
| ChatGPT | OAI-SearchBot | GPTBot | ChatGPT-User fetches pages when a user asks, and robots.txt may not apply to it |
| Perplexity | PerplexityBot | (not used for training, according to its documentation) | Perplexity-User acts at a user’s request |
| Claude | Claude-SearchBot | ClaudeBot | Claude-User acts at a user’s request |
| Google (AI Overviews, AI Mode) | Googlebot | Google-Extended controls use for Gemini and does not affect inclusion in Search | No special markup needed |
Names and behaviours change; the sources at the end of this article are the official pages.
What kinds of page get cited
When we save the sources of answers to buying questions in Spain, the same types of page keep coming up. It is not a published study yet, so take it as a working observation, not as market data:
- comparisons and “best X for Y” lists;
- pricing and “how much does it cost” pages;
- FAQs and “how to choose” guides;
- directories and company listings;
- community threads where someone asks the same thing;
- specialist media and press releases;
- official documentation and institutional pages.
Rarely a home page. Almost never a generic “about us” page.
A citation is not a cause
A page appearing as a source next to an answer means the platform retrieved it and found it useful for that question, on that day. It does not mean the page “caused” a brand’s mention, that the platform always uses it or that it will use it tomorrow. With two rounds a week apart you will see that sources vary too.
What to do with this
- Measure twelve buying questions on two platforms with search and save the cited domains.
- Count which domains repeat. That is your list of the sources that matter in your category.
- For each one, check your presence: an accurate listing, up-to-date details, a real identity.
- For the questions where the cited source is a pricing page or a comparison, check whether your website has that page.
- Let in the search crawlers you care about; decide on the training crawlers separately.