# Perplexity SEO Guide: Check Access, Grade Tactics, Get Cited

> Perplexity SEO guide: a copyable PerplexityBot access check, tactics graded documented, measured or unverified, and September 2026 citation shares.

- URL: https://missiongrowth.io/blog/perplexity-seo
- Published: 2026-07-03 · Updated: 2026-10-04
- Author: Furkan Aktaş, Co-Founder, Mission Growth
- Publisher: Mission Growth. Company facts: https://missiongrowth.io/llms.txt

Perplexity SEO is the work of getting your pages found, fetched and cited when Perplexity answers a question. It is a short chain of gates: the crawler reaches the page, a passage on it answers the query, and your brand appears where Perplexity looks.

The evidence behind each gate is uneven, and that shapes this guide. Perplexity documents the crawler and firewall step, its developer documentation suggests how pages are cut into passages and dated, and Ahrefs measures where citations land.

Weekly refreshes, schema, llms.txt and the popular "results in a few weeks" timelines have no primary evidence at all.

Some teams file this under [answer engine optimization](https://missiongrowth.io/blog/geo-vs-aeo-vs-llmo-vs-aio). That label spans every engine, while this page covers Perplexity AI optimization only, in the order the evidence supports.

In this guide:

- How to get cited by Perplexity AI, starting with which parts of source selection a page owner can change
- What PerplexityBot and Perplexity-User each do to your robots.txt
- A copyable access check for the firewall, the HTML and the logs
- Which domains Perplexity cites most, and what that costs you to chase
- A Perplexity SEO audit that grades each tactic by evidence

## How Perplexity SEO works: the index, the passage and the date

Three levers shape Perplexity visibility: an index that Ahrefs says is Perplexity's own and that tracks Google rank only partly, plus passage extraction and two separate page dates, which Perplexity documents for its developer API.

If you are asking how does Perplexity AI choose sources, only the first lever is a third party's measurement. The other two come from Perplexity's developer documentation, so the step to the consumer product is our inference. Each lever gets its own rule below.

### Perplexity SEO differs from Google SEO: it runs its own index

Ahrefs says Perplexity has its own search index, built from PerplexityBot, and does not draw on Google or Bing's index. Perplexity makes no such statement in the sources we read, so treat it as Ahrefs' account.

That is why a Google ranking is a weak proxy. In Ahrefs' August 2025 study of 15,000 long-tail queries, 28.6% of Perplexity's cited URLs landed in Google's top 10, against about 8% for the other assistants. BrightEdge's April 2024 study had put the overlap at 60%.

The two numbers sit side by side here:

| Study | Date | What it counted | Share in Google's top 10 |
|---|---|---|---|
| BrightEdge | April 2024 | Perplexity citations | 60% |
| Ahrefs | August 2025 | Perplexity cited URLs, 15,000 long-tail queries | 28.6% |

::figure{src="/blog/figures/perplexity-seo-1.svg" alt="Bar chart comparing BrightEdge's 2024 figure of 60% with Ahrefs' 2025 figure of 28.6% for Perplexity citations that rank in Google's top 10." caption="Two studies put the share of Perplexity's cited URLs that rank in Google's top 10 at 60% in 2024 and 28.6% in 2025. By BrightEdge and Ahrefs." width="720" height="182"}

Read the Ahrefs figure two ways. Next to other assistants it is high, which makes Perplexity the most Google-aligned assistant in that sample. Next to 2024 it is low: the earlier figure is 2.1 times the later one.

That leaves 71.4% of cited URLs outside the top 10, so a top 10 rank neither gates nor guarantees a citation. Why the studies differ is our inference: different query sets, more than a year apart, and citations in one study against cited URLs in the other. Neither study isolates the cause.

One correction while we are on BrightEdge. Its "nearly 40% month over month since January" measured referrals from Perplexity to brand sites and says nothing about searches. Don't quote it as search growth, and for current usage figures see [Perplexity AI statistics](https://missiongrowth.io/blog/perplexity-statistics).

### Perplexity's API reads passages, not whole pages

Perplexity's Search API documentation defines its low context setting as "short passages most relevant to the query", and a separate parameter caps the content extracted from each result page. The default is high, which returns detailed content, so the passage setting is a choice the caller makes.

Perplexity also says the API "uses the same search system as the UI with differences in configuration". That supports shared data access. It does not say the consumer answer path extracts the same way.

Here is our inference, labeled as one. If the answer engine also extracts passages conditioned on the query, a section can win a citation on a page whose opening is weak, and an answer buried in the middle is a passage no query retrieves.

The rules that follow cost little either way:

- **Answer under every heading.** The first sentence under a heading answers that heading's question, one claim per sentence.
- **Name the entity.** Write "The pricing page lists the plans" instead of "it lists them", so an extracted passage makes sense alone.
- **Keep answer and support together.** Don't split an answer across a heading boundary.

For example, take a section headed "How much does project management software cost for a small team?"

Before: "Pricing is something every buyer asks about, and the answer depends on many factors that we cover below."

After: "Project management software for a small team usually costs a monthly fee per user, and the price moves with the number of seats and the features you add. The sections below break down each tier."

The second opening survives being lifted out of the page. The first says nothing a query could use. The same reasoning covers headings, lists and short sentences: if the engine returns passages, these are the units it can lift.

### Perplexity's API filters on two dates per page

Perplexity's Search API filters work on a publication date and, separately, a last-updated date. In the API, two independent dates mean a cosmetic re-date can move at most one of them, and a filter on the publication date ignores your refresh entirely.

The documentation never says where either date comes from: page text, markup, headers or crawl time. So showing a visible publication date and a visible last-modified date, and repeating both in structured data, is inference, and no documentation states it as a rule.

Move the updated date only when the substance changed. No primary source gives a refresh cadence, so "update weekly" is unverified, and the audit table later in this guide grades it that way.

## PerplexityBot and Perplexity-User: what each one does

PerplexityBot builds the search index and follows robots.txt, while Perplexity-User fetches a page live for one person's question and generally ignores robots.txt.

Perplexity's crawler documentation separates the two agents like this:

| | PerplexityBot | Perplexity-User |
|---|---|---|
| Purpose | Surfaces and links sites in Perplexity search results; not used to crawl content for AI foundation models | Visits a page to support a user's request |
| robots.txt | Governed by it; Perplexity recommends allowing it | Generally ignores it, because a user requested the fetch |
| IP file | [perplexitybot.json](https://www.perplexity.com/perplexitybot.json) | [perplexity-user.json](https://www.perplexity.com/perplexity-user.json) |

Allowing PerplexityBot is the visibility switch. Blocking it does not stop a live fetch that a user triggers, so a robots.txt block is a poor tool if your goal is to keep Perplexity-User away.

Perplexity also asks you to permit requests from the published IP ranges, which matters in the next section. Robots.txt changes can take up to 24 hours to show in Perplexity's systems, according to its [crawler documentation](https://docs.perplexity.ai/docs/resources/perplexity-crawlers), so wait a day before judging an edit.

The cross-vendor block-or-allow decision for other [AI crawlers](https://missiongrowth.io/blog/gptbot-ai-crawlers) lives in its own guide. Here the job is narrower: get PerplexityBot through every gate.

## Check PerplexityBot access: robots.txt, firewall, HTML and logs

PerplexityBot reaches a page only if robots.txt allows it, the firewall lets its user agent and published IP ranges through, the server returns the text in plain HTML, and the log shows a 200 status.

::figure{src="/blog/figures/perplexity-seo-2.svg" alt="Perplexity SEO access flow chain of four gates PerplexityBot must pass: robots.txt, firewall, plain HTML text and a 200 status in the server log." caption="PerplexityBot reaches a page only if all four gates pass, and a robots.txt allow is only the first." width="720" height="233"}

A robots.txt allow is gate one of four. On a CDN or WAF, read the edge security log as well as the origin log before you touch any content. For the wider crawl and render checks, the [technical GEO checklist](https://missiongrowth.io/blog/technical-geo) goes further.

### Why a robots.txt allow is not enough

Perplexity's documentation says a site behind a web application firewall may need to explicitly allow its bots. Its Cloudflare example sets a rule that combines a User-Agent condition with an IP-address condition against the published IP files.

The pairing matters because a user agent string alone can be spoofed. Two conditions together are harder to fake than either one.

Cloudflare's own moves raise the stakes. It de-listed Perplexity as a verified bot in August 2025, and in July 2025 it changed its default to block AI crawlers. Our inference: on Cloudflare, don't assume a verified bot allow setting covers Perplexity. Add the documented rule, then read the logs.

### What Cloudflare's stealth crawler finding changes

Cloudflare's August 2025 report says Perplexity also crawled with a generic Chrome on macOS user agent from IPs outside its published ranges when its declared crawler was blocked. The common retelling stops there and treats the declared crawlers as the compliant ones.

Cloudflare says otherwise. Both the declared and the undeclared crawlers were accessing content for scraping contrary to the web crawling norms in RFC 9309, so don't describe the declared agents as clean.

The finding also answers a question site owners ask: if I block Perplexity, am I invisible? Cloudflare ran its test on brand-new, unindexed domains. When the stealth crawler was blocked, Perplexity used other data sources, including other websites, and the answers were less specific and lacked details from the original content.

Keep that at its measured scope. It is a result about detail on test domains, and it says nothing about whether a real, indexed brand still gets named.

The decision rule: a block cost detail on Cloudflare's test domains, and a user agent block did not hold against the undeclared crawler. So a block is a network layer decision, made with the firewall, and robots.txt alone won't enforce it.

### Run a short log check

Start with the status codes PerplexityBot and Perplexity-User received:

```bash
grep -E "PerplexityBot|Perplexity-User" access.log | awk '{print $9}' | sort | uniq -c
```

Then compare each hit's IP with the two published IP files, perplexitybot.json and perplexity-user.json. This script prints named bot hits, and also lists browser-like requests (Chrome on macOS) as candidates for review:

```python
import ipaddress, json, re, sys, urllib.request

URLS = ["https://www.perplexity.com/perplexitybot.json",
        "https://www.perplexity.com/perplexity-user.json"]
nets = [ipaddress.ip_network(p["ipv4Prefix"])
        for u in URLS
        for p in json.load(urllib.request.urlopen(u))["prefixes"]
        if "ipv4Prefix" in p]

for line in open(sys.argv[1]):
    ip = line.split()[0]
    named = re.search(r"PerplexityBot|Perplexity-User", line)
    candidate = "Macintosh" in line and "Chrome/" in line and not named
    if named or candidate:
        known = any(ipaddress.ip_address(ip) in n for n in nets)
        status = line.split('"')[2].split()[0]
        print("named" if named else "browser-like", ip,
              "in-published-range" if known else "OUTSIDE-range", status)
```

The script assumes the combined log format, with the client IP first and the status after the request. It reads IPv4 prefixes only. The browser-like lines are candidates and never verdicts, because real Mac visitors match them too.

A CDN changes what the log can tell you. Behind Cloudflare's reverse proxy, your origin server logs a Cloudflare IP address unless you restore the original visitor IP from the CF-Connecting-IP header. Cloudflare documents the server steps, including mod_remoteip for Apache and the real IP module for Nginx.

Run the script on a log with restored client IPs, or it will report every hit as outside the ranges. And a request the firewall blocks at the edge never reaches the origin log, so a 403 that the firewall caused shows in the CDN's security events.

With restored IPs, read the output like this:

- **A 200 status from a published IP:** the gate is open.
- **A 403 while robots.txt allows PerplexityBot:** look at the firewall rule first.
- **No PerplexityBot hits at all:** nobody is fetching, the edge is blocking, or the log is incomplete.
- **A named bot outside the ranges:** a spoof candidate, so check it before you allow it.

JavaScript gets one clause here. Vercel's December 2024 analysis found that none of the major AI crawlers render JavaScript, PerplexityBot included, so text that only appears after scripts run is invisible to it. The full verification steps live in the cross-vendor crawler guide linked earlier.

## What Perplexity cites most, and what to do about it

In Ahrefs' September 2026 snapshot Reddit is Perplexity's most cited source at 21.6% of top 50 citations, ahead of YouTube at 20.8% and Wikipedia at 6.3%.

Reddit overtook YouTube in that snapshot, which moved up one place while YouTube fell one, so "YouTube is the most cited source" is out of date. Ahrefs refreshes the page every month.

The top five domains in the snapshot:

::dataset{key="perplexity-top-cited-domains" name="Perplexity most cited domains by mention share, September 2026"}

| Domain | Mention share | Pages cited |
|---|---|---|
| Reddit | 21.6% | 892,114 |
| YouTube | 20.8% | 1,059,603 |
| Wikipedia | 6.3% | 110,188 |
| Facebook | 4.5% | 158,230 |
| Amazon | 3.6% | 105,853 |

::figure{src="/blog/figures/perplexity-seo-3.svg" alt="Bar chart of the five most cited domains in Perplexity: Reddit 21.6%, YouTube 20.8%, Wikipedia 6.3%, Facebook 4.5% and Amazon 3.6%." caption="Reddit and YouTube lead the top 50 citation pool in Ahrefs' September 2026 snapshot, with Wikipedia a distant third." width="720" height="302"}

The scope matters. Ahrefs defines mention share as a domain's citations as a percentage of the summed citations of the top 50 sources only. Its sample is more than 3.1 million US queries across all topics, so the head leans toward consumer domains such as Healthline, Alibaba, Amazon and Good Housekeeping and says little about B2B queries.

Add up the first three rows and Reddit, YouTube and Wikipedia hold 48.7% of the pool.

Breadth is only one way to be cited. Divide each domain's mention share by its pages cited, then index YouTube to 1.0. The result, citations per cited page, changes the picture:

::dataset{key="perplexity-citations-per-page" name="Perplexity citations per cited page, indexed to YouTube, September 2026"}

| Domain | Mention share | Pages cited | Citations per page (YouTube = 1.0) |
|---|---|---|---|
| YouTube | 20.8% | 1,059,603 | 1.0 |
| Reddit | 21.6% | 892,114 | 1.2 |
| Wikipedia | 6.3% | 110,188 | 2.9 |
| Healthline | 2.6% | 14,312 | 9.3 |
| The New York Times | 1.6% | 6,259 | 13.0 |

::figure{src="/blog/figures/perplexity-seo-4.svg" alt="Bar chart of citations per cited page indexed to YouTube at 1.0: Reddit 1.2, Wikipedia 2.9, Healthline 9.3, The New York Times 13.0." caption="Per cited page, The New York Times is cited about 13 times as often as YouTube and Healthline about 9 times." width="720" height="262"}

YouTube and Reddit win on breadth, with about a million and 892,114 distinct pages. Healthline and The New York Times win on repeat citation of far fewer pages. We computed the index ourselves from Ahrefs' shares and page counts, and the shares carry rounding to one decimal, so read the figures as about.

For a brand, a few reference grade pages cited repeatedly is a different bet from many thread replies. That is our inference, and Ahrefs' data does not say why the pattern exists. Usage and traffic numbers are deliberately left to the statistics page linked above.

### Perplexity SEO and Reddit: what the citation share means

Reddit leads the September 2026 snapshot, but the lead is a share of the top 50 pool across all topics, so it is a reason to check your own prompts and no reason to start posting. Three moves follow from the table:

- **Answer real questions** in the subreddits your buyers read, because Reddit's lead is the largest single block.
- **Publish one clear video** on the question your buyers ask, since YouTube holds nearly as much.
- **Keep your first-party facts consistent** everywhere. On Cloudflare's unindexed test domains, a blocked stealth crawler left Perplexity answering from other websites, so what those sites say is what such an answer can draw on.

### Perplexity Pages and parasite SEO: what the evidence supports

Perplexity's own domain is not among the ten most cited domains in Ahrefs' September snapshot, and Perplexity Pages has no primary evidence as a traffic channel, so treat any "parasite" promise as unverified and don't build a plan on it.

## Which tactics have evidence: a Perplexity SEO audit

An audit sorts each tactic by evidence: Perplexity documents access, its developer API suggests passages and dates, a third party measures the citation mix, and the rest has no primary evidence.

Weekly refreshes, schema markup, llms.txt and topic clusters sit in that last group. The effort column is editorial judgment, and nothing in the table is a ranking factor Perplexity has confirmed.

### How to optimize for Perplexity AI, in order

| Tactic | Evidence grade | Effort | Do it? |
|---|---|---|---|
| Allow PerplexityBot in robots.txt and permit its published IP ranges | Documented by Perplexity | Low | Yes, first |
| Add the firewall allow rule where a WAF sits in front | Documented by Perplexity | Low to medium | Yes, first |
| Serve the text in plain HTML | Measured for PerplexityBot by Vercel, December 2024 | Medium | Yes, first |
| Self-contained passages under each heading | Inference from the Search API documentation | Medium | Yes, next |
| Visible published and updated dates | Inference from the two API date filters | Low | Yes, next |
| Presence on Reddit, YouTube and Wikipedia | Measured citation share, cause not shown | Ongoing | Only after the decision rule below |
| Refresh weekly or every few weeks | Unverified for Perplexity | Ongoing | Optional |
| Schema markup, [llms.txt](https://missiongrowth.io/blog/llms-txt-guide) and topic clusters | Unverified for Perplexity | Low to ongoing | Optional |

In words: do rows one to three first, because they are the gates. Do the passages and dates next. Fund community presence only after the next section's decision rule, and treat the unverified rows as optional.

Two notes on the table. We migrated our own React single-page app to prerendered static HTML for 20 marketing pages because AI crawlers do not execute JavaScript.

On the last row, Perplexity's crawler guidance for site owners names no llms.txt file or schema markup. missiongrowth.io publishes its own llms.txt and llms-full.txt, which is cheap to keep, but don't expect either file to move Perplexity citations.

### Why your brand is invisible in Perplexity despite strong SEO

Strong Google rankings leave a brand invisible in Perplexity when a gate fails silently. Match the symptom to the first check:

- **403 in the logs:** the firewall is blocking, so fix the allow rule.
- **Empty HTML:** the text arrives through JavaScript, so serve it in the first response.
- **Answer buried:** rewrite the opening under each heading so it answers on its own.
- **Competitors cited via Reddit threads:** you are absent from the community layer, which is the case for the decision rule below.

### Measure it: a GA4 segment and a buyer prompt list

Keep measurement simple and hand the tool choice over. In GA4, build a segment or exploration on sessions whose source contains perplexity.ai. Then write a list of real buyer prompts, run them on a fixed cadence in a fresh session, and log the date, whether your domain is cited and which domains are.

To compare tools or go deeper on how to [track Perplexity rankings](https://missiongrowth.io/blog/perplexity-rank-tracking), that guide has the log format. Mission Growth's platform tracks AI citations and visibility for customers.

## Is optimizing for Perplexity worth it?

The access and passage steps are worth doing for any site that wants Perplexity visibility, because a robots.txt change shows in about a day (up to 24 hours), and worth the community presence step only after your own prompts show competitors cited where you are absent.

The split follows what you can verify. Access changes can be checked within a day, while off-site presence is where the measured citation share concentrates and is a distribution commitment, so it should wait for your own prompt evidence.

Here is the rule:

1. **Check your own data.** Look at GA4 for perplexity.ai referral sessions and at your server logs for PerplexityBot.
2. **Run your buyer prompts** and note which domains Perplexity cites.
3. **Match the gap to the fix.** If competitors are cited via Reddit threads or YouTube explainers you are absent from, fund that. If they are cited from their own domains, fix passages and dates first.
4. **Don't size the program from other people's traffic studies,** which range widely. Use your own sessions, and see the [Perplexity usage and referral numbers](https://missiongrowth.io/blog/perplexity-statistics) for the published ranges.

## The Perplexity Publishers Program and Comet Plus

In its August 2025 announcement, Perplexity said it would distribute all Comet Plus revenue to participating publishers, minus a small portion for compute, and that publishers join by emailing publishers@perplexity.ai.

Perplexity Publishers Program payout history, and [where the Comet Plus figures come from](https://missiongrowth.io/blog/perplexity-statistics), are on the statistics page.

This guide stays on Perplexity. For the other engines, read [Gemini SEO](https://missiongrowth.io/blog/gemini-seo) and [ChatGPT SEO](https://missiongrowth.io/blog/chatgpt-seo).

For the cross-engine playbook, see [how to optimize for AI search engines](https://missiongrowth.io/blog/how-to-optimize-for-ai-search-engines).

Perplexity SEO is a short chain of gates with uneven evidence, so work through it in order. Fix access first, write self-contained passages, show honest dates, and fund community presence once your own prompts show you are missing from it. Start today with the log check on your server.

## FAQ

### How long does Perplexity SEO take?

The only documented lag is for robots.txt: Perplexity says its systems can take up to 24 hours to reflect a change. No primary source gives a time to first citation, so be wary of promises of weeks. Run the log check first, because a blocked bot makes every later step irrelevant.

### Does Perplexity use Google's rankings?

Perplexity does not rely on Google's index, according to Ahrefs, which says it runs its own index built from PerplexityBot. In Ahrefs' August 2025 study, 28.6% of its cited URLs ranked in Google's top 10, so a Google rank helps a little and guarantees nothing.

### What is Perplexity SEO?

Perplexity SEO is optimizing so that Perplexity can reach your pages, extract a passage that answers the query and show your brand as a source. It has three gates: access, passage and presence. The broader label, answer engine optimization, covers every engine and not only this one.

### Does PerplexityBot respect robots.txt, and what about Perplexity-User?

PerplexityBot builds the search index, and Perplexity recommends allowing it in robots.txt. Perplexity-User fetches a page because a user asked, so Perplexity says it generally ignores robots.txt rules. Allowing PerplexityBot is the visibility switch, and blocking it does not stop a user triggered fetch.

### Does llms.txt help you get cited by Perplexity?

Nothing in Perplexity's crawler guidance for site owners mentions an llms.txt file, so there is no documented effect. It is cheap to keep, and missiongrowth.io publishes its own, but don't count on it for Perplexity citations. Fix access and passages first.

### Why is my brand invisible in Perplexity when I rank on Google?

A Google rank doesn't prove PerplexityBot can reach you. Common causes are a firewall rule that blocks its user agent, text that only appears after JavaScript runs, an answer buried in the middle of the page, and competitors cited through Reddit threads. The log check shows which gate fails.
