Skip to content
AI & GEO 5 min read

AI & GEO: Optimizing for AI-Powered Search

What AI search assistants actually need from a website — crawler access, indexing and snippet permission — which AI crawlers to allow in robots.txt, what llms.txt and structured data do and do not do, and figures from our study of 4,645 EU websites.

RP
By RankProof
Editorial Team · RankProof

Short answer

To appear in AI search answers, a page first has to be crawlable and indexable by the search systems those answers draw on, and allowed to be shown as a snippet. Google says AI Overviews and AI Mode need no special files or markup; ChatGPT search, Claude and Perplexity use their own crawlers, which robots.txt can allow or block. Beyond access, the same things that make a page good for search make it quotable: a clear answer, facts with sources, and text a machine can read.

Special files needed for Google AI features
None
Controls that limit use
noindex · nosnippet · max-snippet
AI search crawlers
OAI-SearchBot · Claude-SearchBot · PerplexityBot
EU sites blocking no AI crawler
81.7% of 4,645
EU sites with a valid llms.txt
9.4%

AI & GEO: Optimizing for AI-Powered Search

"Generative engine optimisation" (GEO) is the name given to getting a page cited in AI answers — Google's AI Overviews and AI Mode, ChatGPT search, Claude, Perplexity, Copilot. Much of what is sold under that name has no documented basis. This guide separates what the operators document from what is guesswork, and gives figures from our own study of EU websites.

What the operators actually document

Google. Its documentation on AI features says there are no additional requirements to appear in AI Overviews or AI Mode and no special optimisations: a page must be indexed and eligible to be shown in Search with a snippet. The usual Search best practices apply — allow crawling, link pages internally, keep important content in text, keep structured data consistent with the visible page. You do not need AI text files or special markup.

OpenAI, Anthropic, Perplexity. Each runs separate crawlers for separate purposes, and each says its crawlers respect robots.txt:

OperatorSearch / answer crawlerFetches for a user's requestTraining crawler
OpenAIOAI-SearchBotChatGPT-UserGPTBot
AnthropicClaude-SearchBotClaude-UserClaudeBot
PerplexityPerplexityBotPerplexity-User—
GoogleGooglebot (Search, AI Overviews, AI Mode)—Google-Extended (a control token, not a crawler)

The important split: blocking a training crawler does not remove you from AI search, and blocking a search crawler does. If you want to be cited in ChatGPT search, for example, OAI-SearchBot must be allowed; GPTBot can be blocked independently. Google-Extended controls whether content is used for Gemini models and does not affect inclusion in Google Search. Microsoft's Copilot draws on Bing's index, so Bingbot access matters there.

Step 1: access

  1. robots.txt — allow the search and user crawlers you want to be cited by. Rules follow RFC 9309: the most specific matching group applies, and the longest matching path wins.
  2. No accidental blocks at the network level. Bot protection can challenge crawlers that robots.txt allows; check your firewall or CDN logs for blocked requests from the crawlers above.
  3. Indexing — no noindex on pages you want cited.
  4. Snippet permission — nosnippet, data-nosnippet and max-snippet limit what Google may show, including in AI features. Use them deliberately, not by template default.

What EU websites do today

In September 2026 we checked the robots.txt and /llms.txt of popular national websites in all 27 EU countries — up to 200 per country from the Chrome UX Report list for August 2026. Of the 4,645 sites we could measure:

FindingShare of sites
Block no AI crawler at all81.7%
Block at least one training crawler16.5%
Block at least one AI search or user crawler10.1%
Block every AI search and user crawler4.3%
Block GPTBot13.8%
Block OAI-SearchBot6.7%
Publish a valid /llms.txt (of 4,260 checked)9.4%

Two lessons. Most sites block nothing, so access is rarely the problem. And sites that do block tend to block training crawlers more than search crawlers, which matches the split above. Our study "AI crawler blocking on EU websites" has the figures per country and per crawler, and the method.

Step 2: be worth quoting

AI answers quote pages that answer a question clearly and can be trusted. None of this is a secret ranking factor; it is what Google's helpful-content guidance asks for anyway:

  • Answer first. State the answer in the opening lines, then explain. A reader skimming — human or machine — should find it without scrolling.
  • One page, one question. Pages that mix several topics are harder to quote.
  • Facts with sources. Numbers with a date and a link to where they come from; primary sources over summaries.
  • Tables and lists for data that would otherwise be buried in prose.
  • Original information. Your own measurements, examples and experience are what an AI answer cannot get elsewhere — and the reason to cite you.
  • Clear authorship and dates. Who publishes the page, how it was produced, and when it was last changed in substance.

What llms.txt and structured data do — and do not do

llms.txt is a proposal for a Markdown file at /llms.txt that points language models to a site's key pages. It is cheap to publish and harmless, but no major AI search operator documents reading it, and Google says no such file is needed. Treat it as optional; only 9.4% of the EU sites we measured have a valid one.

Structured data (JSON-LD with schema.org types) helps search engines understand entities on a page — an organisation, a product, an article — and must match the visible content. It does not get a page into AI answers by itself.

Measuring

There is no complete report of AI citations. What you can see: in Search Console, clicks from AI Overviews and AI Mode are counted in the Performance report together with web results; visits from ChatGPT, Perplexity, Copilot and similar assistants appear as referrals in your analytics; and server logs show which AI crawlers fetched which pages. Treat any tool that claims a full "AI visibility score" across assistants with caution.

Checking it with RankProof

RankProof's AI readiness check scores only the signals operators document: whether AI search and user crawlers (plus Googlebot and Bingbot) may read the page, whether it is indexable, and whether snippet controls block it. Optional extras — llms.txt, a Markdown version, structured data — are listed but not scored. The robots.txt checker shows which crawlers your rules allow, the robots.txt generator writes rules per crawler group, and the llms.txt generator drafts the file if you decide to publish one.

Ready to rank higher?

Use RankProof's free tools to improve your search rankings and work toward EU accessibility compliance.

Try Free Scan