Perplexity publishes two fetchers, opposite robots.txt rules and a firewall whitelist step, but nothing about how it chooses citations. The one peer-reviewed test on Perplexity found keyword stuffing scored below making no change. On nine Indian page-one domains, Perplexity and ChatGPT split four to four.
You allow the crawler, you publish the answer, and then you wait. Perplexity never tells you why it picked someone else.
What does Perplexity actually publish about getting cited?
Two fetchers, and the difference between them decides whether you are discoverable at all.
| Fetcher | Purpose, in Perplexity's words | robots.txt |
|---|---|---|
| PerplexityBot | "designed to surface and link websites in search results on Perplexity" | Obeyed. Perplexity recommends allowing it |
| Perplexity-User | visits a page "when users ask Perplexity a question" | "generally ignores robots.txt rules" |
PerplexityBot is the one that matters for discovery. Block it and you are not in the pool that gets surfaced. Perplexity-User only fetches a page a user already steered it toward, so it cannot make you findable.
We read every entry in its published documentation index on 27 August. One page is written for a website owner, and it covers crawling and firewall access. None of them states how Perplexity chooses which sources go into an answer.
Is allowing PerplexityBot in robots.txt enough?
No, and this is the step most sites miss. The Perplexity Crawlers page asks for two things, not one.
- Allow PerplexityBot in robots.txt
- Permit requests from its published IP ranges at your firewall
Perplexity devotes a section of that page to Cloudflare and AWS WAF configuration. The instruction is to combine a user-agent match with an IP match, then set the action to Allow.
"A firewall sits in front of robots.txt. It can block a bot your robots.txt has explicitly welcomed, and nothing in Search Console will tell you."
If you run a WAF and have never tested it against a published bot IP list, start there. Perplexity also states changes can take up to 24 hours to appear, so a same-day retest proves nothing.
Not sure whether Perplexity can even reach your site? We check the crawler path, the firewall and how you actually appear across ChatGPT, Gemini and Perplexity.
What has anyone measured on Perplexity itself?
One peer-reviewed study, and it is the only one worth citing. Aggarwal and colleagues ran the KDD 2024 GEO experiments against Perplexity.ai as a live engine, on a 200-query sample.
- 1
Quotation Addition improved Position-Adjusted Word Count by 22% over the baseline
- 2
Statistics Addition improved Subjective Impression by 37%
- 3
Keyword Stuffing scored about 10% below the baseline on Position-Adjusted Word Count, on that same sample
Be careful with the third one. On the paper's other measure, Subjective Impression, Keyword Stuffing went up, from 24.7 to 28.1. It loses on the metric the authors lead with and gains on the softer one.
"The classic SEO move was the only method that went backwards on the measure the study treats as visibility."
That sample used file uploads, because Perplexity does not let a user pin source URLs. So it measures how your content is treated once retrieved, not how it gets retrieved in the first place.
Does Perplexity cite more than ChatGPT?
Not on the set we measured. We took the first nine organic domains on Google India for one commercial query, then read each domain's AI citation counts.
| Domain | DR | Perplexity | ChatGPT |
|---|---|---|---|
| indiamart.com | 87 | 38,732 | 9 |
| oracle.com | 93 | 5,057 | 8,833 |
| frappe.io | 83 | 96 | 97 |
| sagesoftware.co.in | 47 | 21 | 33 |
| matiyas.com | 36 | 11 | 3 |
| jm-origin.com | 22 | 5 | 4 |
| fortunesys.com | 45 | 1 | 0 |
| tabsyst.com | 34 | 1 | 1 |
| udyogsoftware.com | 17 | 0 | 1 |
The split is even, and that is the point.
- Perplexity led on four domains, ChatGPT led on four, one tied
- Perplexity cited 8 of 9 domains, ChatGPT cited 8 of 9
- Neither engine was the harder door to get through
Why one domain decided the answer
Pool the same nine domains and the story inverts.
- All nine: Perplexity 43,924 citations against ChatGPT 8,981, a lead of about 4.9 to 1
- Drop the single largest domain: Perplexity 5,192 against ChatGPT 8,972, and ChatGPT is ahead
The same nine sites, one exclusion, and the winner changes. Any engine comparison built on pooled totals is really a report about its biggest domain.
There is also a problem with the instrument. Across the four named engines on these nine domains, 6 of 36 readings, or 16.7%, reported more distinct cited pages than citation links.
"Distinct cited pages cannot outnumber citation links. We treated those readings as unusable rather than dividing one by the other."
Read all of this as a probe, not a rate:
- One query, one country, one day, nine domains
- Third-party index counts, not our own prompting of either engine
- A zero can mean the index has not covered a site, not that the engine refuses to cite it
Related reading: our breakdown of why a site goes missing from AI answers, what separates a domain every AI engine cites, and where SEO and GEO actually diverge.
See how you show up in ChatGPT, Gemini and Perplexity. Our free AI visibility audit reads your domain across the engines and sends you what it finds. No deck, no pitch.
Key takeaways
- Allow PerplexityBot in robots.txt, then check your firewall against Perplexity's published IP list. The second step is the one sites skip
- Perplexity-User ignores robots.txt by design, so it is not a discovery lever
- Nothing Perplexity publishes explains how it selects sources. Anyone who tells you the ranking factors is guessing
- The only peer-reviewed Perplexity test found quotations and statistics helped. Keyword stuffing lost ground on the study's headline visibility measure and gained on its softer one
- Never trust a pooled engine comparison. One domain moved ours by a factor of five
Frequently asked
Does blocking PerplexityBot remove me from Perplexity?
It removes you from what PerplexityBot surfaces, which is the discovery path. Perplexity-User can still fetch your page when a user points it there, because that fetcher generally ignores robots.txt.
Is there a Perplexity ranking factor list?
Not a published one. Perplexity's documentation covers its crawlers and its APIs. No page in its documentation index sets out how sources are selected for an answer.
What actually improves Perplexity citations?
On the only peer-reviewed test, n=200, adding quotations and adding statistics. Both beat the baseline on Perplexity itself. Keyword stuffing scored about 10% below the baseline on Position-Adjusted Word Count, the measure that paper leads with.
Can I check whether Perplexity cites my site?
Third-party AI citation indexes report per-engine counts, but treat them carefully. Our own read found internally inconsistent numbers in 6 of 36 readings, and a zero can mean index coverage rather than absence.
Sources
Adith Krishnan
Co-Founder & COO, Kula Digital
8+ years building marketing that is measured in revenue. Runs strategy, AI search visibility, and education marketing at Kula. Written from the studio in Coimbatore.
