Kula Digital

LLM SEO: what Google, OpenAI and Perplexity publish

We spent a morning reading the crawler documentation instead of the commentary. The interesting part was not what the platforms optimise for. It was what they exempt.

A

Adith Krishnan

Co-Founder & COO, Kula Digital

Published
2026-08-12
Read time
7 min read
Updated
2026-08-12
Category
GEO
TL;DRThe short version

LLM SEO is not a separate discipline with its own index. Google says its AI answers retrieve from the ordinary Search index. The real gap is crawler control: all three platforms we read publish user-triggered fetchers that generally ignore robots.txt, and on Google that now covers two AI products.

There is no LLM index to optimise for. That single fact settles most of what gets sold under the name. What it does not settle is who is allowed to fetch your pages, and that turns out to be the part almost nobody has checked.

What does "LLM SEO" actually mean?

The term is used for at least four different jobs. They carry different costs, different controls and different odds of paying back, so quoting for all four as one line item is how budgets go missing.

  • Getting cited inside an AI answer on Google. Google's own systems, retrieving from the ordinary Search index
  • Getting surfaced in ChatGPT search. A separate crawler with its own robots.txt control
  • Getting surfaced in Perplexity. Separate again, and controlled separately again
  • Being used in model training. A different question entirely, and one you may want to refuse

Only the first three are marketing work. Google addressed the naming directly in July 2026, and the answer it gave is the reason this post exists.

Does Google run a separate index for AI answers?

No. Google's guide to optimizing for generative AI features says its generative features are rooted in its core Search ranking and quality systems, and describes retrieval-augmented generation as relying on those systems to pull pages from the Search index. Grounding is the other name Google gives the same technique.

"From Google's perspective there is nothing to optimise separately, because the thing doing the retrieving is the search index you already have."

The same guide names the terms head on. It says AEO and GEO both describe work aimed at AI search visibility, and that optimizing for generative AI search is optimizing for the search experience, and thus still SEO. That is Google describing its own product, so weigh it accordingly, but no other platform has published anything this direct about the naming.

How many crawlers are you actually managing?

More than the one most site owners think about. We read the current crawler documentation published by each platform on 12 August 2026 and counted the separately addressable agents.

PlatformSearch agentOther agents you can address
Google SearchGooglebotGoogle-Extended, for Gemini apps and Vertex AI
OpenAIOAI-SearchBotGPTBot, OAI-AdsBot, ChatGPT-User
PerplexityPerplexityBotPerplexity-User
Counted from the official crawler documentation of each platform, read in full on 12 August 2026. Google-Extended is listed separately because Google states it governs Gemini apps and Vertex AI rather than Search. Three platforms, not a survey.

OpenAI and Perplexity both split search from training and say each setting works independently. So blocking the training crawler does not remove you from the search answers, and a great many robots.txt files were written as though it does.

Which of these ignore your robots.txt?

Here is the finding we did not expect. Every one of the three platforms publishes a class of fetcher that is exempt, and all three describe it in nearly the same words.

PlatformExempt fetchersWhat the documentation says
Google9 listed user-triggered fetchersGenerally ignore robots.txt rules
OpenAIChatGPT-Userrobots.txt rules may not apply
PerplexityPerplexity-UserGenerally ignores robots.txt rules
Counted from each platform's published crawler documentation on 12 August 2026. Google's page states its list is not exhaustive. Wording is condensed by us, the positions are each platform's own.

Google's list of user-triggered fetchers puts it plainly: because the fetch was requested by a user, these fetchers generally ignore robots.txt rules. Perplexity says the same of Perplexity-User, and OpenAI says robots.txt rules may not apply to ChatGPT-User.

The part worth acting on

Two of the nine on Google's list are AI products: Gemini Notebook and Google-Agent, the latter described as used by agents to navigate the web and act on user request. If your plan for AI visibility was a robots.txt edit, it does not reach these.

Want this checked against your own domain? Here is what a free AI visibility audit covers.

See how you show up in ChatGPT, Gemini and Perplexity. Free AI visibility audit.

What does Google say you can skip?

Google's July 2026 guide includes a mythbusting section, which is unusual and worth reading in the original. It names five things you do not need to do.

What people sellWhat Google's guide says
llms.txt and AI-specific filesGoogle Search ignores them
Chunking content for AINo requirement, no ideal page length
Rewriting copy for AI systemsNot needed, synonyms are understood
Buying mentions across the webInauthentic mentions are not as helpful as they seem
Special schema for AINot required, no special markup exists
Summarised from Google's guide to optimizing for generative AI features, last updated 10 July 2026. The wording is ours, the positions are Google's.

Note the limit of this. It is Google telling you what Google ignores, and it does not bind OpenAI or Perplexity, neither of whom publishes an equivalent list.

What actually decides whether you appear?

The requirement Google does state is ordinary and unglamorous: a page must be indexed and eligible to be shown with a snippet. Same bar as classic search, which means the failures are the old failures.

What we foundSites
Homepage not fully readable without JavaScript5 of 30
No meta description at all12 of 30
Description broken, generic or belonging to another page4 of 30
Kula crawlability audit. 30 institution homepages fetched as raw HTML with no JavaScript execution, August 2026. Of the 5 unreadable, 2 returned nothing at all and 3 returned only fragments.

A page a crawler cannot read is absent from every one of these systems at once, which is why we start there rather than with the AI-specific tactics. The full diagnostic sits in the two reasons a site is absent from AI answers.

Does any of this matter for a local business yet?

Less than the volume of noise suggests. In our pull of 160 India marketing queries, not one of the 30 commercial queries carrying a city name returned an AI Overview.

That zero is not a zero everywhere. At n=30 it is consistent with a true rate of up to 9.5% at 95% confidence, so read it as uncommon rather than impossible. The query-level method sits in our analysis of which query types AI Overviews take.

The crawler work still pays, because it is the same work classic search has always needed and it costs a morning. The AI-specific spend is what we would hold, at least until a city-qualified search starts returning an AI answer in your own category and your own town.

What this cannot tell you

  • Three platforms is not the market. We read Google, OpenAI and Perplexity, because those are the platforms whose documentation sits on our approved fetch list. Anthropic and Microsoft are unread here, and their absence is a limit of our method rather than a finding about them
  • Documentation is a statement of intent, not observed behaviour. We did not test whether any crawler obeys what its own page says, and the exemptions above are the places where that gap would matter most
  • Google's own list says it is not exhaustive. Nine is what is published, not what exists
  • The audit sample is one category. 30 institution homepages, one country, one month. Your sector may fail differently
  • Platform documentation changes without notice. Every count here is dated 12 August 2026 and should be rechecked rather than trusted in six months

Key takeaways

  • There is no separate LLM index, and Google states its generative features retrieve from the ordinary Search index
  • Google names AEO and GEO directly and calls the work SEO, in a guide last updated 10 July 2026
  • All 3 platforms we read publish user-triggered fetchers that their own documentation says generally ignore robots.txt
  • 2 of the 9 fetchers on Google's exempt list are AI products, Gemini Notebook and Google-Agent
  • Blocking GPTBot does not remove you from ChatGPT search answers, because OpenAI controls the two separately
  • 5 of 30 homepages we audited were not fully readable without JavaScript, which fails every one of these systems at once

Frequently asked

Is LLM SEO a real thing?

The work is real, the separate discipline is not. Google's guide states that optimising for generative AI search is optimising for the search experience and is still SEO. What is genuinely new is crawler management, because OpenAI and Perplexity publish agents that Googlebot rules do not cover, and those need deciding on separately.

Do I need an llms.txt file?

Not for Google. Its guide says you do not need to create machine readable files, AI text files, markup or Markdown to appear in Google Search including its generative features, and that Google Search ignores them. It also says keeping one for other systems will neither harm nor help your Google visibility. Other platforms publish no equivalent statement either way.

If I block GPTBot, do I disappear from ChatGPT?

Not from its search answers. OpenAI documents GPTBot as the training crawler and OAI-SearchBot as the search one, and states the settings are independent, so a site can allow search while disallowing training. OpenAI says it is sites opted out of OAI-SearchBot that will not be shown in ChatGPT search answers.

Can I actually stop an AI product from reading my site?

Not entirely with robots.txt alone. All three platforms exempt user-triggered fetches, and Google's own documentation says its user-triggered fetchers generally ignore robots.txt. For Feedfetcher, Google's guidance is to serve an error status to that user agent instead. Anything stronger than robots.txt is a server-side decision, not a file in your root.

Does structured data help me appear in AI answers?

Google says it is not required and that no special schema markup exists for generative AI search, while still recommending it for rich results in ordinary search. We keep it on client sites for that reason rather than for AI visibility, and we would treat any agency selling schema as an AI visibility tactic as selling something Google has publicly called unnecessary.

Should a college or clinic in Tamil Nadu care about this yet?

Cautiously. In our 160 query pull, none of the 30 commercial queries carrying a city name returned an AI Overview, so the immediate exposure looks low for locally searched businesses. The part worth doing now is the crawlability work, because it pays in classic search regardless of what the AI surfaces do next.

Sources

  1. 1.Optimizing your website for generative AI features, Google Search Central
  2. 2.List of Google user-triggered fetchers, Google Crawling Infrastructure
  3. 3.Overview of OpenAI Crawlers, OpenAI
  4. 4.Perplexity Crawlers, Perplexity
A

Adith Krishnan

Co-Founder & COO, Kula Digital

8+ years building marketing that is measured in revenue. Runs strategy, AI search visibility, and education marketing at Kula. Written from the studio in Coimbatore.

More on GEO

Related reading

All articles