In this article 4
Short answer
AI with search can only cite what a crawler has been able to read and index. If your pages are not in the index, load their content only with JavaScript, return errors, are duplicated by badly set canonicals or take too long, no content or source work will help. It is the least glamorous part of GEO and the one that most often explains a “we do not show up”.
The five-minute check
- Search Google for site:yourdomain.com. Do your important pages appear? With the right titles?
- Open Search Console and look at Page indexing. How many pages are excluded, and why?
- Open a key page with JavaScript disabled (or look at the page source). Is the main text there?
- Open yourdomain.com/robots.txt. Are you blocking something you should not? AI crawlers are decided one by one.
- Measure speed on mobile. Does the content load in under a few seconds?
If any of these fails, stop here and fix it before going further.
The errors we see most
Content only in JavaScript. Google can process JavaScript, but in separate phases and with more complexity; other crawlers, less so. If the essential text is not in the initial HTML, for many systems it does not exist. Prerender or statically generate whatever you want cited.
Pages excluded by a canonical. A canonical that points to another URL tells the search engine “index the other one”. If the other one does not answer the question, you have hidden your best page.
Inherited noindex. From the development phase, from a plugin, from a template. Check the pages that matter one by one.
Redirect chains and soft 404s. A page that returns 200 with a “not found” message confuses the index. A 404 should be a 404.
CDN blocks. Anti-bot rules that return 403 to legitimate crawlers. You only see them in the logs.
Duplicates from parameters or languages. Versions with and without a trailing slash, with and without www, in two languages without hreflang. Every duplicate splits signals.
An outdated sitemap. With URLs that no longer exist, or without the new ones. A sitemap does not index anything by itself, but it helps discovery.
What to look at on a small services website
- One self-referencing canonical URL per indexable page.
- A unique, descriptive title and description.
- An H1 that names the service and the market.
- Main text in HTML, not in images or PDFs.
- A sitemap submitted in Search Console and Bing Webmaster Tools.
- A robots.txt that allows public pages and search crawlers.
- Real 404s for routes that do not exist.
- No non-essential cookies before consent where the law requires it (it is not crawling, but you will be asked about it anyway).
When this is “SEO” and when it is “GEO”
It is the same thing. These fundamentals serve classic search and AI search alike, because systems with search retrieve from the same kind of index. That is why a GEO-only agency also checks crawling and indexing: not as an SEO service, but as a requirement for the measurement to make sense.