Perplexity is the only major AI platform that has published how its retrieval system works, and the details contradict a fair amount of standard AEO advice.
It is also the platform where citation matters most. Every Perplexity answer carries visible sources, users arrive expecting to click, and the company has reported roughly 45 million active users and 780 million monthly queries with growth above 370% year on year, raising at a $20 billion valuation (TechCrunch).
Here is what its own engineering write-up says about how sources get picked.
Perplexity's published research on architecting an AI-first search API and its crawler documentation, rather than inference from observed results. Where we are extrapolating, we say so.
Does Perplexity use Google's index?
No. It runs its own.
Perplexity describes an exabyte-scale index and crawling apparatus that scales horizontally with both index size and query volume, with crawler and indexing fleets comprising tens of thousands of CPUs and hundreds of terabytes of RAM (Perplexity Research). Its crawlers are documented separately, including PerplexityBot (Perplexity).
That independence shows up in the data. Ahrefs found 80% of sources cited by AI search platforms do not appear in Google at all, with only 12% matching Google's top results (Ahrefs).
The practical consequence is direct: your Google ranking is not an input. If PerplexityBot cannot crawl you, or crawls you rarely, you are not a candidate regardless of how you rank.
How does Perplexity decide what to index?
This is the part most worth reading twice, because it names a preference almost nobody optimises for.
Perplexity says it uses learned models and tuned heuristics to decide which documents to keep hot, prioritising documents from authoritative domains, documents covering undercovered topics, and other categories with outsize potential to enrich the index (Perplexity Research).
Authoritative domains is expected. Undercovered topics is not, and it is a genuine strategic opening. Perplexity is explicitly biased toward content on subjects its index is thin on. The fiftieth article on "what is AEO" adds nothing to an index already saturated with them. A properly researched piece on a narrow question nobody has covered well is exactly what the system is built to reward.
For most businesses this is better news than the alternative. Competing on authority against established publishers is slow. Competing on coverage of your own specific niche is winnable this quarter.
Perplexity states it prioritises indexing documents on undercovered topics. Broad, heavily-written subjects are the hardest place to earn a citation. The specific, unglamorous question your customers actually ask is the easiest.
How does Perplexity decide when to re-crawl?
By predicting how often a page changes.
Perplexity's training objective is calibrated to both the importance and the likely update frequency of specific URLs, and it uses machine learning to schedule when indexing happens (Perplexity Research).
So a page that changes meaningfully on a predictable rhythm earns more frequent crawls than one that never moves. This aligns with what is measured across the wider ecosystem: AI platforms cite content 25.7% fresher than what ranks organically (Ahrefs).
The lesson is not to churn pages for the sake of it. It is that a genuinely maintained page compounds, because being crawled more often means being a live candidate more often.
Why do citations quote one paragraph rather than a whole page?
Because Perplexity does not score pages. It scores parts of pages.
Its indexing and retrieval infrastructure divides documents into fine-grained units, which are individually surfaced and scored against the original query parameters (Perplexity Research).
This single design decision explains most of what makes AEO different from SEO. Your page is not competing as a page. Each self-contained section competes on its own. A 4,000-word guide is not one entry in the contest, it is thirty, and the weak ones do not drag down the strong ones.
It also explains why thin, well-structured pages sometimes beat comprehensive ones. If your answer to a specific question is buried in paragraph nine and depends on paragraphs one through eight for context, the fragment that gets scored is incoherent on its own. A competitor who answers it in three self-contained sentences under a clear heading wins that fragment.
| Traditional SEO instinct | What fragment-level scoring rewards |
|---|---|
| One page per keyword | One self-contained answer per section |
| Build to a conclusion | Answer first, then expand |
| Context carried across the page | Each section standing alone |
| Comprehensive coverage of a broad topic | Complete coverage of a narrow one |
| Publish more | Maintain what exists |
What actually gets you cited by Perplexity?
Five things follow from the architecture, ordered by how directly the published material supports them.
Be crawlable by PerplexityBot. Check your robots.txt, your firewall and your bot-management rules. Plenty of sites block AI crawlers by default through a WAF preset and have no idea. This is the one failure that makes everything else irrelevant.
Write self-contained sections. Each heading a question, each opening paragraph a complete answer to it. This is not stylistic advice, it is fitting the unit of retrieval.
Go narrow before broad. The index is biased toward undercovered topics. Your specific, awkward, low-volume questions are a better bet than another general explainer.
Maintain a real update rhythm. Crawl scheduling is predicted from update frequency, and cited content runs 25.7% fresher than organic results.
Build authority signals off-site. Ahrefs found branded web mentions had the strongest correlation with AI visibility of any factor measured, at 0.664 (Ahrefs). Perplexity's stated preference for authoritative domains points the same way.
Is Perplexity worth the effort given its size?
On raw traffic, Perplexity is the smallest of the major platforms. On traffic quality, it is arguably the best.
Perplexity is citation-first by design. Sources are visible in every answer and users expect to click through, which makes its referral traffic behave more like traditional search traffic than ChatGPT's does. Set against that, ChatGPT accounts for more than 80% of all AI referral traffic (Ahrefs), and Google still sends around 345 times more traffic than every AI platform combined (Ahrefs).
The honest position: do not build a Perplexity-specific programme. Build content that works at fragment level, keep it current, stay crawlable, and Perplexity is one of several engines that rewards it. The overlap is the point, as we argue in our AI search market share breakdown. For platform-specific tactics, our guide to ranking in Perplexity goes deeper.
The most common reason businesses are invisible to Perplexity
It is almost never content quality. It is access.
Perplexity crawls with its own bots, and a great many sites block AI crawlers without anyone having decided to. Managed WAF rule sets on Cloudflare, AWS and similar platforms increasingly ship with AI bot categories enabled by default, and enterprise bot-management products often treat any non-search-engine crawler as suspicious. The result is a site that ranks well in Google, converts well, and cannot be cited by Perplexity because PerplexityBot has never successfully fetched a page.
Review your robots.txt, your CDN or WAF bot rules, and your server logs for PerplexityBot requests. If you see requests being served 403s, every other optimisation on this page is irrelevant until that is fixed.
Blocking is a legitimate choice for some publishers, particularly those whose business is selling access to their content. It should just be a choice rather than a default someone inherited from a security preset.
How Perplexity differs from ChatGPT and Gemini
The three engines reach conclusions differently enough to be worth distinguishing.
| Perplexity | ChatGPT | Gemini | |
|---|---|---|---|
| Index | Own exabyte-scale index | Web search plus training data | Google Search and Maps grounding |
| Citations | Visible on every answer | Sometimes shown | Inline, attached to text segments |
| Local data | Web sources | Web sources and listings | Google Business Profile via Maps |
| Stated bias | Authoritative domains and undercovered topics | Not published | Not published in detail |
| Scoring unit | Fine-grained document fragments | Passage level | Text segment level |
The commonality is more useful than the differences. All three score fragments rather than pages, all three favour fresher content, and all three respond to the same off-site credibility signals. Yext found 86% of AI citations come from brand-managed sources (Yext), which holds regardless of which engine is asking.
Build for fragment-level retrieval once, keep it current, stay crawlable, and you have covered the material differences between all of them.
Find out whether Perplexity cites you
Book a free AI Visibility Audit. We check Perplexity, ChatGPT, Google AI Overviews, Gemini and Claude, and show you which sources they use instead of you.
Book Your AI Visibility AuditFrequently Asked Questions
Does Perplexity use Google's search index?
No. Perplexity operates its own exabyte-scale index and crawling fleet (Perplexity Research). Ahrefs found 80% of sources cited by AI platforms do not appear in Google at all, and only 12% match Google's top results.
What is PerplexityBot and should I allow it?
It is Perplexity's crawler, documented in their crawler resources. If you block it, whether deliberately or through a firewall preset, your content cannot be indexed and cannot be cited. Allowing it is the prerequisite for everything else.
Why does Perplexity cite a single paragraph from my page?
Because it divides documents into fine-grained units and scores each individually against the query (Perplexity Research). Sections compete separately, so a self-contained answer under a clear heading is more citable than the same information spread across a page.
Does Perplexity favour large, established sites?
It prioritises authoritative domains, but it also explicitly prioritises documents on undercovered topics (Perplexity Research). That second preference is what gives smaller sites a genuine route in, provided they cover something the index is thin on.
How often should I update content for Perplexity?
Often enough to establish a pattern, because crawl scheduling is calibrated to predicted update frequency (Perplexity Research). Meaningful revisions to important pages beat cosmetic edits across many pages.

Ashur Homa
Built and scaled a digital brand to $100M+ in sales with zero ad spend. Has helped businesses generate millions through AI go-to-market strategy. Leads growth at Omni Eclipse.
Connect on LinkedIn