Perplexity SEO Guide: Check Access, Grade Tactics, Get Cited
Perplexity SEO guide: a copyable PerplexityBot access check, tactics graded documented, measured or unverified, and September 2026 citation shares.
By Furkan AktaşPublished Updated

On this page
Perplexity SEO is the work of getting your pages found, fetched and cited when Perplexity answers a question. It is a short chain of gates: the crawler reaches the page, a passage on it answers the query, and your brand appears where Perplexity looks.
The evidence behind each gate is uneven, and that shapes this guide. Perplexity documents the crawler and firewall step, its developer documentation suggests how pages are cut into passages and dated, and Ahrefs measures where citations land.
Weekly refreshes, schema, llms.txt and the popular "results in a few weeks" timelines have no primary evidence at all.
Some teams file this under answer engine optimization. That label spans every engine, while this page covers Perplexity AI optimization only, in the order the evidence supports.
In this guide:
- How to get cited by Perplexity AI, starting with which parts of source selection a page owner can change
- What PerplexityBot and Perplexity-User each do to your robots.txt
- A copyable access check for the firewall, the HTML and the logs
- Which domains Perplexity cites most, and what that costs you to chase
- A Perplexity SEO audit that grades each tactic by evidence
How Perplexity SEO works: the index, the passage and the date
Three levers shape Perplexity visibility: an index that Ahrefs says is Perplexity's own and that tracks Google rank only partly, plus passage extraction and two separate page dates, which Perplexity documents for its developer API.
If you are asking how does Perplexity AI choose sources, only the first lever is a third party's measurement. The other two come from Perplexity's developer documentation, so the step to the consumer product is our inference. Each lever gets its own rule below.
Perplexity SEO differs from Google SEO: it runs its own index
Ahrefs says Perplexity has its own search index, built from PerplexityBot, and does not draw on Google or Bing's index. Perplexity makes no such statement in the sources we read, so treat it as Ahrefs' account.
That is why a Google ranking is a weak proxy. In Ahrefs' August 2025 study of 15,000 long-tail queries, 28.6% of Perplexity's cited URLs landed in Google's top 10, against about 8% for the other assistants. BrightEdge's April 2024 study had put the overlap at 60%.
The two numbers sit side by side here:
Read the Ahrefs figure two ways. Next to other assistants it is high, which makes Perplexity the most Google-aligned assistant in that sample. Next to 2024 it is low: the earlier figure is 2.1 times the later one.
That leaves 71.4% of cited URLs outside the top 10, so a top 10 rank neither gates nor guarantees a citation. Why the studies differ is our inference: different query sets, more than a year apart, and citations in one study against cited URLs in the other. Neither study isolates the cause.
One correction while we are on BrightEdge. Its "nearly 40% month over month since January" measured referrals from Perplexity to brand sites and says nothing about searches. Don't quote it as search growth, and for current usage figures see Perplexity AI statistics.
Perplexity's API reads passages, not whole pages
Perplexity's Search API documentation defines its low context setting as "short passages most relevant to the query", and a separate parameter caps the content extracted from each result page. The default is high, which returns detailed content, so the passage setting is a choice the caller makes.
Perplexity also says the API "uses the same search system as the UI with differences in configuration". That supports shared data access. It does not say the consumer answer path extracts the same way.
Here is our inference, labeled as one. If the answer engine also extracts passages conditioned on the query, a section can win a citation on a page whose opening is weak, and an answer buried in the middle is a passage no query retrieves.
The rules that follow cost little either way:
- Answer under every heading. The first sentence under a heading answers that heading's question, one claim per sentence.
- Name the entity. Write "The pricing page lists the plans" instead of "it lists them", so an extracted passage makes sense alone.
- Keep answer and support together. Don't split an answer across a heading boundary.
For example, take a section headed "How much does project management software cost for a small team?"
Before: "Pricing is something every buyer asks about, and the answer depends on many factors that we cover below."
After: "Project management software for a small team usually costs a monthly fee per user, and the price moves with the number of seats and the features you add. The sections below break down each tier."
The second opening survives being lifted out of the page. The first says nothing a query could use. The same reasoning covers headings, lists and short sentences: if the engine returns passages, these are the units it can lift.
Perplexity's API filters on two dates per page
Perplexity's Search API filters work on a publication date and, separately, a last-updated date. In the API, two independent dates mean a cosmetic re-date can move at most one of them, and a filter on the publication date ignores your refresh entirely.
The documentation never says where either date comes from: page text, markup, headers or crawl time. So showing a visible publication date and a visible last-modified date, and repeating both in structured data, is inference, and no documentation states it as a rule.
Move the updated date only when the substance changed. No primary source gives a refresh cadence, so "update weekly" is unverified, and the audit table later in this guide grades it that way.
PerplexityBot and Perplexity-User: what each one does
PerplexityBot builds the search index and follows robots.txt, while Perplexity-User fetches a page live for one person's question and generally ignores robots.txt.
Perplexity's crawler documentation separates the two agents like this:
Allowing PerplexityBot is the visibility switch. Blocking it does not stop a live fetch that a user triggers, so a robots.txt block is a poor tool if your goal is to keep Perplexity-User away.
Perplexity also asks you to permit requests from the published IP ranges, which matters in the next section. Robots.txt changes can take up to 24 hours to show in Perplexity's systems, according to its crawler documentation, so wait a day before judging an edit.
The cross-vendor block-or-allow decision for other AI crawlers lives in its own guide. Here the job is narrower: get PerplexityBot through every gate.
Check PerplexityBot access: robots.txt, firewall, HTML and logs
PerplexityBot reaches a page only if robots.txt allows it, the firewall lets its user agent and published IP ranges through, the server returns the text in plain HTML, and the log shows a 200 status.
A robots.txt allow is gate one of four. On a CDN or WAF, read the edge security log as well as the origin log before you touch any content. For the wider crawl and render checks, the technical GEO checklist goes further.
Why a robots.txt allow is not enough
Perplexity's documentation says a site behind a web application firewall may need to explicitly allow its bots. Its Cloudflare example sets a rule that combines a User-Agent condition with an IP-address condition against the published IP files.
The pairing matters because a user agent string alone can be spoofed. Two conditions together are harder to fake than either one.
Cloudflare's own moves raise the stakes. It de-listed Perplexity as a verified bot in August 2025, and in July 2025 it changed its default to block AI crawlers. Our inference: on Cloudflare, don't assume a verified bot allow setting covers Perplexity. Add the documented rule, then read the logs.
What Cloudflare's stealth crawler finding changes
Cloudflare's August 2025 report says Perplexity also crawled with a generic Chrome on macOS user agent from IPs outside its published ranges when its declared crawler was blocked. The common retelling stops there and treats the declared crawlers as the compliant ones.
Cloudflare says otherwise. Both the declared and the undeclared crawlers were accessing content for scraping contrary to the web crawling norms in RFC 9309, so don't describe the declared agents as clean.
The finding also answers a question site owners ask: if I block Perplexity, am I invisible? Cloudflare ran its test on brand-new, unindexed domains. When the stealth crawler was blocked, Perplexity used other data sources, including other websites, and the answers were less specific and lacked details from the original content.
Keep that at its measured scope. It is a result about detail on test domains, and it says nothing about whether a real, indexed brand still gets named.
The decision rule: a block cost detail on Cloudflare's test domains, and a user agent block did not hold against the undeclared crawler. So a block is a network layer decision, made with the firewall, and robots.txt alone won't enforce it.
Run a short log check
Start with the status codes PerplexityBot and Perplexity-User received:
grep -E "PerplexityBot|Perplexity-User" access.log | awk '{print $9}' | sort | uniq -c
Then compare each hit's IP with the two published IP files, perplexitybot.json and perplexity-user.json. This script prints named bot hits, and also lists browser-like requests (Chrome on macOS) as candidates for review:
import ipaddress, json, re, sys, urllib.request
URLS = ["https://www.perplexity.com/perplexitybot.json",
"https://www.perplexity.com/perplexity-user.json"]
nets = [ipaddress.ip_network(p["ipv4Prefix"])
for u in URLS
for p in json.load(urllib.request.urlopen(u))["prefixes"]
if "ipv4Prefix" in p]
for line in open(sys.argv[1]):
ip = line.split()[0]
named = re.search(r"PerplexityBot|Perplexity-User", line)
candidate = "Macintosh" in line and "Chrome/" in line and not named
if named or candidate:
known = any(ipaddress.ip_address(ip) in n for n in nets)
status = line.split('"')[2].split()[0]
print("named" if named else "browser-like", ip,
"in-published-range" if known else "OUTSIDE-range", status)
The script assumes the combined log format, with the client IP first and the status after the request. It reads IPv4 prefixes only. The browser-like lines are candidates and never verdicts, because real Mac visitors match them too.
A CDN changes what the log can tell you. Behind Cloudflare's reverse proxy, your origin server logs a Cloudflare IP address unless you restore the original visitor IP from the CF-Connecting-IP header. Cloudflare documents the server steps, including mod_remoteip for Apache and the real IP module for Nginx.
Run the script on a log with restored client IPs, or it will report every hit as outside the ranges. And a request the firewall blocks at the edge never reaches the origin log, so a 403 that the firewall caused shows in the CDN's security events.
With restored IPs, read the output like this:
- A 200 status from a published IP: the gate is open.
- A 403 while robots.txt allows PerplexityBot: look at the firewall rule first.
- No PerplexityBot hits at all: nobody is fetching, the edge is blocking, or the log is incomplete.
- A named bot outside the ranges: a spoof candidate, so check it before you allow it.
JavaScript gets one clause here. Vercel's December 2024 analysis found that none of the major AI crawlers render JavaScript, PerplexityBot included, so text that only appears after scripts run is invisible to it. The full verification steps live in the cross-vendor crawler guide linked earlier.
What Perplexity cites most, and what to do about it
In Ahrefs' September 2026 snapshot Reddit is Perplexity's most cited source at 21.6% of top 50 citations, ahead of YouTube at 20.8% and Wikipedia at 6.3%.
Reddit overtook YouTube in that snapshot, which moved up one place while YouTube fell one, so "YouTube is the most cited source" is out of date. Ahrefs refreshes the page every month.
The top five domains in the snapshot:
Download CSV (CC BY 4.0)
The scope matters. Ahrefs defines mention share as a domain's citations as a percentage of the summed citations of the top 50 sources only. Its sample is more than 3.1 million US queries across all topics, so the head leans toward consumer domains such as Healthline, Alibaba, Amazon and Good Housekeeping and says little about B2B queries.
Add up the first three rows and Reddit, YouTube and Wikipedia hold 48.7% of the pool.
Breadth is only one way to be cited. Divide each domain's mention share by its pages cited, then index YouTube to 1.0. The result, citations per cited page, changes the picture:
Download CSV (CC BY 4.0)
YouTube and Reddit win on breadth, with about a million and 892,114 distinct pages. Healthline and The New York Times win on repeat citation of far fewer pages. We computed the index ourselves from Ahrefs' shares and page counts, and the shares carry rounding to one decimal, so read the figures as about.
For a brand, a few reference grade pages cited repeatedly is a different bet from many thread replies. That is our inference, and Ahrefs' data does not say why the pattern exists. Usage and traffic numbers are deliberately left to the statistics page linked above.
Perplexity SEO and Reddit: what the citation share means
Reddit leads the September 2026 snapshot, but the lead is a share of the top 50 pool across all topics, so it is a reason to check your own prompts and no reason to start posting. Three moves follow from the table:
- Answer real questions in the subreddits your buyers read, because Reddit's lead is the largest single block.
- Publish one clear video on the question your buyers ask, since YouTube holds nearly as much.
- Keep your first-party facts consistent everywhere. On Cloudflare's unindexed test domains, a blocked stealth crawler left Perplexity answering from other websites, so what those sites say is what such an answer can draw on.
Perplexity Pages and parasite SEO: what the evidence supports
Perplexity's own domain is not among the ten most cited domains in Ahrefs' September snapshot, and Perplexity Pages has no primary evidence as a traffic channel, so treat any "parasite" promise as unverified and don't build a plan on it.
Which tactics have evidence: a Perplexity SEO audit
An audit sorts each tactic by evidence: Perplexity documents access, its developer API suggests passages and dates, a third party measures the citation mix, and the rest has no primary evidence.
Weekly refreshes, schema markup, llms.txt and topic clusters sit in that last group. The effort column is editorial judgment, and nothing in the table is a ranking factor Perplexity has confirmed.
How to optimize for Perplexity AI, in order
In words: do rows one to three first, because they are the gates. Do the passages and dates next. Fund community presence only after the next section's decision rule, and treat the unverified rows as optional.
Two notes on the table. We migrated our own React single-page app to prerendered static HTML for 20 marketing pages because AI crawlers do not execute JavaScript.
On the last row, Perplexity's crawler guidance for site owners names no llms.txt file or schema markup. missiongrowth.io publishes its own llms.txt and llms-full.txt, which is cheap to keep, but don't expect either file to move Perplexity citations.
Why your brand is invisible in Perplexity despite strong SEO
Strong Google rankings leave a brand invisible in Perplexity when a gate fails silently. Match the symptom to the first check:
- 403 in the logs: the firewall is blocking, so fix the allow rule.
- Empty HTML: the text arrives through JavaScript, so serve it in the first response.
- Answer buried: rewrite the opening under each heading so it answers on its own.
- Competitors cited via Reddit threads: you are absent from the community layer, which is the case for the decision rule below.
Measure it: a GA4 segment and a buyer prompt list
Keep measurement simple and hand the tool choice over. In GA4, build a segment or exploration on sessions whose source contains perplexity.ai. Then write a list of real buyer prompts, run them on a fixed cadence in a fresh session, and log the date, whether your domain is cited and which domains are.
To compare tools or go deeper on how to track Perplexity rankings, that guide has the log format. Mission Growth's platform tracks AI citations and visibility for customers.
Is optimizing for Perplexity worth it?
The access and passage steps are worth doing for any site that wants Perplexity visibility, because a robots.txt change shows in about a day (up to 24 hours), and worth the community presence step only after your own prompts show competitors cited where you are absent.
The split follows what you can verify. Access changes can be checked within a day, while off-site presence is where the measured citation share concentrates and is a distribution commitment, so it should wait for your own prompt evidence.
Here is the rule:
- Check your own data. Look at GA4 for perplexity.ai referral sessions and at your server logs for PerplexityBot.
- Run your buyer prompts and note which domains Perplexity cites.
- Match the gap to the fix. If competitors are cited via Reddit threads or YouTube explainers you are absent from, fund that. If they are cited from their own domains, fix passages and dates first.
- Don't size the program from other people's traffic studies, which range widely. Use your own sessions, and see the Perplexity usage and referral numbers for the published ranges.
The Perplexity Publishers Program and Comet Plus
In its August 2025 announcement, Perplexity said it would distribute all Comet Plus revenue to participating publishers, minus a small portion for compute, and that publishers join by emailing publishers@perplexity.ai.
Perplexity Publishers Program payout history, and where the Comet Plus figures come from, are on the statistics page.
This guide stays on Perplexity. For the other engines, read Gemini SEO and ChatGPT SEO.
For the cross-engine playbook, see how to optimize for AI search engines.
Perplexity SEO is a short chain of gates with uneven evidence, so work through it in order. Fix access first, write self-contained passages, show honest dates, and fund community presence once your own prompts show you are missing from it. Start today with the log check on your server.
Frequently asked questions
How long does Perplexity SEO take?
The only documented lag is for robots.txt: Perplexity says its systems can take up to 24 hours to reflect a change. No primary source gives a time to first citation, so be wary of promises of weeks. Run the log check first, because a blocked bot makes every later step irrelevant.
Does Perplexity use Google's rankings?
Perplexity does not rely on Google's index, according to Ahrefs, which says it runs its own index built from PerplexityBot. In Ahrefs' August 2025 study, 28.6% of its cited URLs ranked in Google's top 10, so a Google rank helps a little and guarantees nothing.
What is Perplexity SEO?
Perplexity SEO is optimizing so that Perplexity can reach your pages, extract a passage that answers the query and show your brand as a source. It has three gates: access, passage and presence. The broader label, answer engine optimization, covers every engine and not only this one.
Does PerplexityBot respect robots.txt, and what about Perplexity-User?
PerplexityBot builds the search index, and Perplexity recommends allowing it in robots.txt. Perplexity-User fetches a page because a user asked, so Perplexity says it generally ignores robots.txt rules. Allowing PerplexityBot is the visibility switch, and blocking it does not stop a user triggered fetch.
Does llms.txt help you get cited by Perplexity?
Nothing in Perplexity's crawler guidance for site owners mentions an llms.txt file, so there is no documented effect. It is cheap to keep, and missiongrowth.io publishes its own, but don't count on it for Perplexity citations. Fix access and passages first.
Why is my brand invisible in Perplexity when I rank on Google?
A Google rank doesn't prove PerplexityBot can reach you. Common causes are a firewall rule that blocks its user agent, text that only appears after JavaScript runs, an answer buried in the middle of the page, and competitors cited through Reddit threads. The log check shows which gate fails.
Cite this page
Aktaş, F. (2026, October 4). Perplexity SEO Guide: Check Access, Grade Tactics, Get Cited. Mission Growth. https://missiongrowth.io/blog/perplexity-seo
Download the data: 2 tables as CSV (CC BY 4.0)
Figures we made for this post are free to reuse under CC BY 4.0 with credit to Mission Growth.
Get Mission Growth highlighted in your Google results.


