A technical reference on AI search visibility. This site sells nothing, takes no engagements and endorses no products. Consulting enquiries are handled separately at hartzer.com.

Hartzer.it.com logoHartzer.it.comAI search visibility reference
Abstract nested ring illustration representing Perplexity
AI Search SurfacePlatform

Perplexity

Perplexity documents its two crawlers with unusual precision, and defines nothing about the word it leans on hardest: authoritative.

UnsupportedPerplexity's own documentation contradicts the standard advice for blocking it, and says nothing about how sources are chosen.

An answer engine, not a search engine with a summary bolted on

Perplexity AI is an answer engine: a search product whose default output is written prose rather than a list of results. A query triggers a live search of the web, the system reads the pages it retrieves, and it returns an answer carrying numbered inline citations, each one a link, with a strip of sources beside or above the text. Perplexity's own definition, in its help center, is that "An answer engine is a tool designed to give you direct, detailed answers to your questions." The company was founded in August 2022 by Aravind Srinivas, Denis Yarats, Johnny Ho and Andy Konwinski, and the product launched publicly on 7 December 2022.

The distinction from ChatGPT matters more than it sounds. In ChatGPT, search is one behavior inside a general assistant. In Perplexity, retrieval with attribution is the product. That makes it the surface where citation behavior is most visible to an outsider, and it is also why its crawling practices have drawn more scrutiny, more publisher anger and more legal challenge than any comparable operator's.

Two structural facts follow. Perplexity has a stronger incentive than its competitors to link out, because links are the interface. And it has a weaker position from which to refuse publishers, because the material it displays is visibly theirs.

Two crawlers, two rule sets, and the mistake almost everyone makes

Perplexity documents two fetching agents and gives them different rules. PerplexityBot is the indexing crawler: in Perplexity's words it is "designed to surface and link websites in search results on Perplexity" and "is not used to crawl content for AI foundation models." Perplexity recommends allowing it. Perplexity-User is the on-demand fetcher: it "supports user actions within Perplexity," visiting a page during a user's question in order to answer it and link to it. Then comes the sentence that matters: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules."

So a robots.txt disallow aimed at Perplexity governs the index, not the answer. A page blocked to PerplexityBot can still be fetched and cited when somebody's question sends Perplexity-User to it. Publishers who block and then see themselves cited are not watching a malfunction; they are watching documented behavior, written down by the operator, in public.

This is not a contested claim or a practitioner observation. It is on Perplexity's own bots page, in the second half, past the part most people stop reading. The belief that one robots.txt line controls the whole product survives because every other article about crawler control treats robots.txt as a single switch, and because the first half of Perplexity's page reads exactly like the pages those articles describe.

Perplexity also publishes the full user-agent strings and JSON files of IP ranges for both agents, and advises combining user-agent matching with IP verification. Take the advice. A user-agent string is a claim, not an identity.

The Cloudflare dispute: an allegation, a denial, and no replication

On 4 August 2025 Cloudflare published findings that Perplexity was using undeclared crawlers to reach content on sites that had blocked it. The stated method was direct: register brand-new domains, unknown to anyone, with a robots.txt disallowing all automated access; ask Perplexity about them; see whether the content comes back. Cloudflare reported that it did, describes an undeclared agent presenting itself as Chrome on macOS, using addresses outside Perplexity's published ranges and rotating between networks at three to six million requests a day. Cloudflare removed Perplexity from its verified-bot list and shipped blocking signatures. The Cloudflare write-up sets out the test in full.

Perplexity replied the next day that Cloudflare had confused it with unrelated traffic from BrowserBase, a third-party cloud browser service, put its own volume below 45,000 daily requests, and called the analysis technically flawed.

Report both. Do not report either as settled. No neutral third party has published a replication of Cloudflare's test, and that replication is precisely the evidence that would resolve it. It is also worth recording that this was the second such report, not the first: in June 2024 Wired and the developer Robb Knight described the same class of behavior — undisclosed addresses and spoofed user-agent strings fetching pages that robots.txt disallowed.

As of 29 August 2026 no Cloudflare post reversing the delisting could be found. Treat the delisting as standing, and check before you rely on it.

Authoritative is asserted, never defined

Perplexity's help center, last updated 1 May 2026, says the product "searches the internet, gathering information from authoritative sources like articles, websites, and journals." That sentence is the whole of the published guidance on source selection. It does not say whose index is searched, how many pages are retrieved, how candidates are scored, how many survive into an answer, or what makes a source authoritative in operational terms.

This is a documentation silence, and it should be named as one rather than worked around. Every "how to get cited in Perplexity" checklist in circulation is inference from observed citation sets — somebody ran prompts, looked at what came back, and generalized. That method is legitimate and some of the resulting advice is probably right. What it is not is documented. Generative engine optimization, meaning the practice of trying to influence what an answer engine cites, has no published criteria to work against here, and any page presenting its recommendations as Perplexity's requirements has invented a source.

The honest form of the advice is conditional: this is what practitioners observe, here is the sample it came from, and it could stop being true without notice or announcement. Anyone unwilling to write that sentence is selling certainty that does not exist.

Being cited pays only inside a program

Perplexity is the only major answer engine that has published a mechanism paying publishers for citation as such. The Publishers' Program launched in July 2024, sharing advertising revenue with publishers whose content is cited; early partners included Time, Der Spiegel, the Texas Tribune, The Independent, Adweek, Gannett and Lee Enterprises. Comet Plus, announced 25 August 2025 and launched 2 October 2025, shares subscription revenue and allocates it, in Perplexity's words, "based on three types of internet traffic: human visits, search citations, and agent actions." Seven launch partners were named: CNN, Condé Nast, Fortune, the Los Angeles Times, the Washington Post, Le Monde and Le Figaro.

Two corrections to how this usually gets reported. First, the widely quoted 80% split to publishers comes from Press Gazette's reporting of 2 October 2025, not from Perplexity, whose own wording is that it distributes the revenue minus a small portion for compute costs. Attribute the number to the journalism, not to the company. Second, and more consequentially: being cited pays nothing unless you are in a program. There is no payment to an ordinary site that happens to appear in an answer, and no published route by which one arrives.

Closest to classic search results, and still mostly outside them

The Ahrefs citation overlap study of 11 August 2025 — Louise Linehan with Xibeijia Guan — ran 15,000 long-tail queries across Google, Bing, ChatGPT, Gemini, Copilot and Perplexity, extracted the citations, and compared them against Google's results. Perplexity had the highest overlap of the six: 28.6% of its citations appeared in Google's top 10 for the same query, against 8.6% for Gemini, 8.2% for Copilot, 8.0% for ChatGPT in-text citations and 6.1% for ChatGPT references. Across all six engines only 12% of AI-cited links appeared in Google's top 10.

Ahrefs attributes Perplexity's higher figure to its being built to cite and running its own index. The practical reading is that conventional ranking work transfers to Perplexity better than to any other surface covered here — and that it still explains under a third of what Perplexity cites. Roughly seven in ten Perplexity citations come from outside Google's top 10, which is the opposite of the conclusion usually drawn from "Perplexity is the most Google-like."

Perplexity's referral position depends on which denominator you use, and the two available answers disagree in rank order. By share of AI referral traffic reaching websites it was 7.23%, third, on SE Ranking's panel of 101,574 sites covering January 2025 to April 2026. By share of traffic to generative AI websites — a usage measure, not a referral one — Similarweb put it at 1.3% in May 2026, seventh. Perplexity punches above its usage share on referrals, which is what a citation-forward interface should do. Neither figure predicts anything about your site.

Measurement rests on a convention nobody promised

Perplexity documents no UTM parameter, no referral convention and no measurement guidance of any kind. What practitioners rely on is that clicks arrive with a perplexity.ai referrer and land in analytics as referral traffic. That works. It is also an observed convention rather than a documented one, and it can change without notice, without announcement, and without anything appearing in the help center afterwards. This is the exact inverse of OpenAI's position, where the tagging parameter is written down and can be quoted back.

Nor is there any impression reporting. No publisher console, no citation counts, nothing outside whatever a Comet Plus partner receives privately — and Perplexity publishes nothing about what partners see either. So a publisher can count clicks and cannot count citations, which means the visible number is a floor of unknown depth.

The publisher lawsuits belong in this picture as context rather than as guidance. News Corp, through Dow Jones and the New York Post, sued in October 2024; Nikkei and Asahi Shimbun sued in Japan on 26 August 2025 seeking ¥2.2 billion each and alleging explicitly that robots.txt was ignored and that summaries carried false attributions; Reddit has also sued, and the BBC, Forbes and Wired have complained publicly. As of the reporting available on 29 August 2026 none of it was resolved. It does not tell you how to be cited. It tells you that the terms on which Perplexity uses publisher content are being decided somewhere other than in its documentation.

Frequently asked questions

Does blocking PerplexityBot in robots.txt keep me out of Perplexity's answers?

No, and Perplexity says so itself. PerplexityBot governs the search index. Perplexity-User is the agent that fetches a page during a user's question, and Perplexity's bots documentation states it "generally ignores robots.txt rules" because the fetch was user-requested. Blocking the crawler removes you from the index while leaving user-triggered fetches untouched.

How do I identify Perplexity referral traffic?

By the referrer host perplexity.ai, which shows up as referral traffic in standard analytics. That is the only method available, because Perplexity documents no UTM parameter and no referral convention at all. It works today and rests on nothing published, so build the segment, label it as an assumption, and re-check it whenever your Perplexity numbers move sharply.

Does Perplexity pay publishers when it cites them?

Only participants in its programs. The Publishers' Program from July 2024 shares advertising revenue; Comet Plus, from October 2025, shares subscription revenue allocated across human visits, search citations and agent actions. A site that is simply cited receives nothing. The frequently quoted 80% split comes from Press Gazette's reporting, not from Perplexity, whose own wording is vaguer.

Was Perplexity proven to be evading robots.txt?

Not proven. Cloudflare published a stated method and its findings on 4 August 2025 and removed Perplexity from its verified-bot list; Perplexity denied it the next day, attributing the traffic to BrowserBase. No neutral third party has replicated Cloudflare's test. Report an allegation with a method and a denial with a specific alternative explanation, and do not report a conclusion.

Do I need to rank in Google to be cited by Perplexity?

It helps more here than anywhere else and still is not sufficient. Ahrefs found in August 2025 that 28.6% of Perplexity citations appeared in Google's top 10 for the same query, the highest of six engines tested. The other side of that number is that roughly seven citations in ten come from pages outside Google's top 10.

Top