Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Cloudflare says it found traffic it attributed to Perplexity reaching websites after Perplexity’s declared crawlers were blocked, sometimes posing as an ordinary Chrome browser. Perplexity’s current policy says its official crawler respects robots.txt. The central unresolved question is whether the traffic Cloudflare identified was in fact operated by Perplexity; the public account is a detailed technical allegation, not an independent finding or court ruling.

What Cloudflare said it found

In an investigation published August 4, 2025, Cloudflare alleged that Perplexity used undeclared crawlers after website owners blocked its named crawlers. Cloudflare identified those declared agents as PerplexityBot and Perplexity-User. It said some customers had also blocked Perplexity-related IP ranges, restricted access to robots.txt, or applied web application firewall (WAF) rules.

Cloudflare said it then tested the behavior on newly purchased domains that it described as undiscoverable and not indexed by search engines. The test sites had restrictive robots.txt rules and additional WAF controls. Cloudflare queried Perplexity about the domains and said the service returned detailed information about their contents. It interpreted this as evidence that another route was retrieving the pages despite the restrictions. Cloudflare’s account of the investigation sets out its methodology and conclusions.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cloudflare said the second traffic pattern used a generic Chrome-on-macOS user-agent rather than identifying itself as Perplexity:

#1 Best Overall
Wall Climbing Spider, Remote Control Robot Toy with LED Eyes for Kids 3+
  • Gravity-Defying Wall & Ceiling Crawling Spider Toy – Ultimate Wall Climbing Spider & Bug Toy: Watch this advanced robotic pet defy gravity! Using powerful suction technology, this wall-crawling spider toy smoothly scales smooth walls, glass, and even ceilings—just like a real spider. Perfect for thrilling races, imaginative adventures, or hilarious pranks, it’s the ultimate interactive wall climbing toy and a top pick for fans of robotic pets, reptile-themed toys, and STEM play. A standout alternative to remote control spider style crawlers.
  • lightweight body design & Safe Materials: Lightweight design lets it cling to walls and minimizes fall impact; high-toughness plastic ensures unbeatable shatter resistance and mimics Spider skin’s bouncy texture for doubled play fun. Responsive controls suit all skill levels—ideal for solo play or friend competitions.Plus, constructed with non-toxic materials, it ensures safe fun for toddlers and older kids alike, totally eliminating parents’ worries over toy safety.
  • 360° Stunts & Lifelike Movements, Super Cool & Eye-Catching Design: Take full command of your RC spider! Pull off thrilling 360° spins with its flexible body and realistic crawl — a perfect blend of RC snake excitement and wall gecko/lizard charm. It’s sure to captivate and spark curiosity! Boasting vibrant colors, 8 flexible multi-jointed legs, and glowing LED eyes that light up in the dark. It encourages imaginative bug-themed adventures and exploration.Safe for ages 3+, it’s an exciting addition to any collection of scary toys, prank toys, or robotic animal gifts.
  • Rechargeable & Easy-Use 2.4GHz Remote Control Spider – Long-Lasting Fun for Kids: Enjoy eco-friendly, hassle-free play! The USB-rechargeable battery (cable included) delivers up to 35 minutes of continuous ground play or 18 minutes of wall-climbing action. The 2.4GHz anti-interference remote (requires 2x AA batteries, not included) ensures stable control up to 115ft, supports multi-player races without signal overlap, and makes it an ideal RC toy for indoor and outdoor group fun—great for birthdays, holidays, or creative kids' electronics.
  • The Perfect Gift for Kids – Great for Holidays, STEM Play & Creative Fun: The ultimate surprise for any occasion! Packaged in a gift-ready box, this wall crawler is a hit for birthdays, Christmas, Halloween,New Year's Day,Valentine's Day,April Fools' Day,Easter, Children’s Day, Back-to-School, or as a fun April Fools’ gag gift. It encourages STEM interest, imaginative play, and endless entertainment. An unforgettable gift for boys and girls who love Spiderman toys, action figures, remote control reptiles, and interactive robot pets.
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7)
AppleWebKit/537.36 (KHTML, like Gecko)
Chrome/124.0.0.0 Safari/537.36

Cloudflare further reported that this traffic came from multiple IP addresses outside Perplexity’s published ranges and changed among autonomous systems (ASNs)—network operators that announce groups of IP addresses. It said the activity appeared across tens of thousands of domains.

Declared and alleged undeclared traffic, according to Cloudflare

Declared crawler traffic Alleged stealth traffic
Identified as PerplexityBot or Perplexity-User Used a generic Chrome/macOS browser identity
Associated with published crawler guidance and IP ranges Cloudflare said some requests came from IPs outside those ranges and from changing ASNs
Reported by Cloudflare at about 20–25 million requests per day Reported by Cloudflare at about 3–6 million requests per day

Those volumes are Cloudflare’s estimates, not independently audited counts. Likewise, the table describes Cloudflare’s classification and attribution; it does not establish that Perplexity operated every request in either group.

Why Cloudflare called it “stealth”

The allegation is broader than a crawler choosing an unfamiliar name. Cloudflare described a combination of an undeclared identity, a browser-like user-agent, IP and ASN rotation, and a fallback pattern it said appeared after named agents were blocked. Together, those signs can make automated requests harder for a site to distinguish from ordinary browsing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

None of those clues alone proves who controlled the traffic. User-agent strings can be copied. Cloud browsers, proxies, distributed hosting, security tools, and third-party data providers can also produce browser-like requests or traffic from changing networks. Attribution is stronger when multiple observations align: the timing after a block, requested URLs, query-specific behavior, content returned in answers, and evidence from the receiving application or infrastructure.

Cloudflare’s newly created, non-indexed test domains matter because they were designed to reduce the chance that the answers merely reflected information already available through a search index. That makes Cloudflare’s direct-access interpretation more compelling, but it does not by itself identify the operator behind every request or rule out all third-party explanations.

Perplexity’s published position

Perplexity’s crawler documentation recommends allowing PerplexityBot and its published IP ranges if a publisher wants content to appear in Perplexity search results. Its help center, updated July 16, 2026, says the official PerplexityBot respects robots.txt and will not index full or partial page text when a site disallows it. The same page says a blocked page may still yield a domain, headline, and brief factual summary; that is distinct from indexing the page’s full text.

Perplexity also says a previously available feature through which users could submit blocked URLs for summaries has been disabled, and that third-party crawlers used to build its index are expected to respect robots.txt, particularly for news publishers. It says content allowed into its search index is not used to pre-train foundation models, noting that Perplexity does not build foundation models. These are the company’s published statements, not independent verification of crawler behavior. Perplexity’s current robots.txt explanation and crawler documentation describe its stated policy.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It is useful to distinguish four categories: PerplexityBot, the declared search crawler; Perplexity-User, a declared user-action crawler named in Cloudflare’s report; third-party crawlers Perplexity says it may use for indexing; and the generic-browser traffic Cloudflare attributed to Perplexity. A policy for the first category does not, by itself, resolve the identity or responsibility for the fourth. Nor does a current policy prove or disprove what happened in 2025.

What robots.txt does—and what it cannot do

robots.txt is a machine-readable set of instructions for crawlers. A site can use it to express which paths compliant crawlers should or should not request. Google’s documentation describes it as crawler guidance and references RFC 9309, the Robots Exclusion Protocol standard.

It is not a password, firewall, or access-control system. A public URL remains technically reachable unless the site imposes an actual barrier. A crawler can ignore a disallow rule; robots.txt alone cannot stop the connection. Do not put confidential material behind a robots.txt rule and assume it is protected.

Ignoring a robots.txt rule may disregard a site’s stated preference, but the file alone does not settle whether access was illegal. Legal questions can turn on jurisdiction, authorization, contract terms, technical barriers, copyright, and the specific conduct. Treat robots.txt as one part of a publishing policy, not as a universal legal test.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cloudflare’s comparison with ChatGPT-User

In the same investigation, Cloudflare said it ran comparable tests involving ChatGPT-User. It reported that the crawler retrieved robots.txt, stopped when access was disallowed, stopped after receiving a block page, and did not continue with follow-up crawls under other user-agents or third-party bots. That is Cloudflare’s account of its tests—not a universal finding about all OpenAI traffic, every ChatGPT feature, or every crawling situation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the public evidence establishes—and what remains unresolved

The public material establishes that Cloudflare published a detailed account of its test design, the traffic patterns it said it observed, its estimated volumes, and its attribution to Perplexity. The test domains and reported traffic signatures are relevant evidence. But Cloudflare is the source for both the observations and the conclusion that Perplexity was responsible; the available account is not a court finding or a separately reproduced forensic investigation.

The key open question is therefore attribution: was the alleged traffic operated directly by Perplexity, by a contractor or infrastructure provider, or by another intermediary? Perplexity’s statements about its official crawler and third-party expectations give its current position, but do not independently answer that question about the traffic Cloudflare reported in 2025.

Cloudflare also has a commercial interest in bot-management and AI-crawler controls. It sells tools to detect, challenge, and block automated traffic while reporting on that traffic and developing publisher controls. That context does not invalidate the technical findings, but readers should keep the distinction between Cloudflare’s measurements and its interpretation in view.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What website owners can do

If you want to tell compliant Perplexity crawlers not to fetch your site, robots.txt can express that preference:

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

This only applies to clients that identify themselves and follow the file. For stronger enforcement, combine clear crawler directives with controls at the network and application layers:

  • Use WAF rules, rate limits, and bot-management scoring. Apply challenges or blocks to suspicious traffic rather than relying on a user-agent string alone.
  • Protect private material with authentication and origin controls. robots.txt is not a substitute for access control.
  • Log and compare requests. Record user-agent, source IP, ASN, requested paths, timestamps, headers, and response status. These signals can support an investigation but may not reveal the ultimate operator.
  • Use IP and ASN restrictions cautiously. Published ranges can change, and third-party infrastructure may not use them. Broad network blocks can affect legitimate visitors.
  • Consider a canary page for testing. A unique, non-sensitive marker on a non-indexed test page can help detect retrieval, but an answer containing the marker is an indicator to investigate, not conclusive attribution by itself.
  • Review what blocked responses reveal. Titles, metadata, error text, or a block page may still disclose information even when the main page is denied.

A practical investigation can start by checking logs for requests after a crawler block, then comparing timing, paths, headers, IPs, and ASNs with the provider’s published guidance. If you test with a unique marker, use a clean test domain, preserve timestamps and logs, avoid sensitive content, and repeat the test before drawing conclusions. A CDN or bot-management service may classify traffic probabilistically; a score or fingerprint is not necessarily proof of ownership.

There are trade-offs. Blocking only known IP ranges can miss third-party infrastructure. Blocking generic Chrome traffic risks stopping real visitors. Blocking every AI crawler may reduce discovery or referrals you want to keep. Decide separately whether to permit search indexing, user-triggered browsing, model-training crawlers, or other categories rather than treating all automated traffic as equivalent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The wider publisher dispute

This dispute reflects a broader shift from conventional search crawling to AI answer engines, retrieval-augmented generation, and user-triggered browsing. Publishers are asking what counts as consent, how a crawler should identify its operator, and what value—referrals, licensing revenue, or other benefits—should flow back when content is used to answer questions.

Cloudflare’s later analysis describes a “crawl-to-click gap”: AI crawlers can generate substantial request volume while sending relatively few referral visits. That framing comes from Cloudflare’s research, but it captures why crawler access has become a commercial as well as a technical issue. Cloudflare’s analysis of AI crawler activity and referrals explores the issue.

Cloudflare has introduced managed robots.txt and AI-crawler controls, including a control intended to block recognized AI scrapers, and has described AI Crawl Control and pay-per-crawl as ways to manage or potentially monetize access. These tools can simplify policy management, but they do not turn robots.txt into a security boundary or guarantee that every undeclared client will be identified. Cloudflare’s managed robots.txt announcement, its AI crawler blocking announcement, and its AI Crawl Control overview explain those offerings.

For publishers, the most durable approach is to make the policy explicit, use technical controls proportionate to the risk, retain evidence, and distinguish crawler categories. For readers evaluating the accusation, the most accurate summary is that Cloudflare reported a serious pattern and attributed it to Perplexity, while the public evidence does not independently establish every step of that attribution.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.