Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Reddit says it set a digital trap for Perplexity—and that the company’s answer engine reproduced a uniquely marked post that should have been discoverable only through Google. That is compelling evidence for Reddit’s theory that Perplexity, or a supplier working for it, obtained Reddit material through indirectly scraped search results. But it is not a court finding that Perplexity personally operated the scraper, violated copyright, or broke the law.
The dispute is now a federal lawsuit over scraping, search indexes, AI retrieval, robots.txt, and who gets to profit from online content.
What Reddit says happened
On October 22, 2025, Reddit sued Perplexity AI, SerpApi, Oxylabs UAB, and AWMProxy in federal court in New York. In its complaint, Reddit alleged that the defendants obtained Reddit content indirectly through Google search-result pages after Reddit tried to restrict unauthorized automated collection.
Reddit described the alleged arrangement as “data laundering.” Its theory was that scraping and proxy companies collected search-result pages containing Reddit text, links, images, and videos, then supplied that information to customers including Perplexity. Perplexity could consequently access representations of Reddit content without making the same direct requests to Reddit that its defenses were designed to block.
#1 Best Overall
Reddit also filed a first amended complaint on February 6, 2026. The allegations described below remain allegations in litigation, not an adjudicated finding.
How the alleged Reddit trap worked
Reddit’s most important evidence was a controlled test post—essentially a digital equivalent of a marked banknote.
- Reddit created a post containing an unusual identifier, described as a hexadecimal string.
- The post was configured so Google could crawl or index it.
- Reddit said the material was not otherwise discoverable through normal public web searches or ordinary Reddit discovery.
- Reddit queried Perplexity using the uncommon identifier.
- According to the complaint, Perplexity reproduced the unusual content within hours.
Reddit inferred that Perplexity, or a vendor acting on its behalf, had obtained the information by scraping Google’s search-result pages. The company’s argument is straightforward: if the post was visible to Google but not normally available elsewhere, and Perplexity quickly returned its distinctive contents, an intermediary may have harvested Google’s representation of the post and fed it into Perplexity’s system.
That makes the test significant, but it does not answer every technical question. The result alone does not establish:
- which company performed the alleged scraping;
- whether Perplexity instructed or knowingly used that scraper;
- whether the information was stored, licensed, cached, or retrieved in real time;
- whether another technical route could have produced the answer; or
- whether the conduct violated a particular law or contract.
What “nearly three billion pages” actually means
Reddit alleged that SerpApi, Oxylabs, and AWMProxy collectively accessed nearly three billion Google search-engine-results pages containing Reddit material during a two-week period in July 2025.
That figure should not be described as three billion stolen Reddit posts or three billion unique articles. It refers to alleged automated accesses to search-result pages. A single page request may contain snippets, links, media references, or repeated material, and the complaint does not make the number equivalent to the number of distinct Reddit works copied.
Reddit also alleged that citations to Reddit content in Perplexity answers increased fortyfold. That figure comes from Reddit’s complaint and contemporary reporting, including Futurism’s account; it is not an independently adjudicated measurement.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Why scrape Google instead of Reddit?
Reddit’s theory is that Google’s index provided an indirect route around Reddit’s defenses.
Reddit can block, rate-limit, or challenge direct automated requests to its own systems. Google, however, had already crawled and indexed portions of Reddit. A scraper targeting Google results could potentially collect Reddit snippets, URLs, images, and related text without making identical direct requests to Reddit.
That distinction is the heart of the case. The accusation is not simply that Perplexity summarized a publicly visible webpage. It is that companies deliberately obtained content through another service to bypass technical or contractual restrictions, then used the resulting data commercially.
Perplexity’s response: retrieval is not model training
Perplexity denied the core accusation in a public response reproduced on Reddit. It said it is an application-layer answer engine rather than a company training foundation models on Reddit content. Perplexity characterized its product as summarizing Reddit discussions and providing citations, and said Reddit was seeking leverage in negotiations over data licensing.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →That response highlights a distinction often lost in coverage:
- Model training involves using data to develop or refine a model’s parameters.
- Live retrieval involves finding or receiving information when a user asks a question.
- Answer generation involves using retrieved information to produce a response, often with citations.
Perplexity’s claim that it does not train foundation models on Reddit content would address one category of use. It would not, by itself, resolve Reddit’s separate allegation that Reddit material was obtained and used in Perplexity’s live answer product. Conversely, evidence of retrieval would not automatically prove that Reddit content was used to train a foundation model.
What role did SerpApi, Oxylabs, and AWMProxy allegedly play?
Reddit’s complaint alleged that:
- SerpApi provides programmatic search-engine scraping or search-result data services.
- Oxylabs provides proxy and web-data collection infrastructure.
- AWMProxy was described by Reddit as a former Russian botnet-related operation.
Reddit alleged that these companies harvested Google results containing Reddit material and made the data available to customers, including Perplexity. Those descriptions come from Reddit’s pleading. Being named as a defendant does not establish that a company committed the alleged conduct, that it acted for Perplexity, or that Perplexity knew how a supplier obtained information.
How Cloudflare’s crawler report fits in
The Reddit lawsuit followed a separate controversy involving Perplexity’s web-crawling practices. In an August 4, 2025 report, Cloudflare said its tests observed Perplexity using both declared and undeclared crawlers.
Cloudflare reported that, after its test domains blocked automated access through robots.txt and web-application-firewall rules, it observed another crawler using a generic browser user agent and IP addresses outside Perplexity’s published range. Cloudflare interpreted that activity as an attempt to evade no-crawl directives.
This provides context, but it is not proof that the Cloudflare tests and Reddit’s test post involved the same infrastructure. Cloudflare reported its own observations; Reddit presented a separate theory involving Google search-result pages.
Why robots.txt does not settle the case
robots.txt is a convention websites use to communicate crawler preferences. It is useful, but it is not a universal security barrier, copyright license, or complete legal regime.
Rank #4
Ignoring a robots rule could become relevant evidence in a dispute about intent, access controls, contract, circumvention, or unfair conduct. But whether it creates legal liability depends on the specific claim, the site’s terms, the technical measures involved, the jurisdiction, and what the defendant actually did.
Publishers therefore face a practical problem: blocking a named crawler may not block requests using different user agents, rotating IP addresses, proxy providers, search indexes, caches, or downstream data suppliers. A robots file communicates a preference; it does not necessarily prevent indirect access.
What legal questions remain unresolved?
Reddit’s allegations potentially implicate several legal theories:
- Copyright infringement: whether protected Reddit material was copied, displayed, or commercially exploited without permission.
- Circumvention: whether technical protections were deliberately bypassed.
- Contract: whether website terms or other agreements restricted the alleged collection or use.
- Interference with computer systems: whether automated requests imposed unauthorized burdens or bypassed access controls.
- Unjust enrichment and unfair competition: whether defendants benefited unfairly from Reddit’s content or infrastructure.
- Supplier responsibility: whether Perplexity can be held responsible for a vendor’s conduct, and whether the evidence connects the vendor’s actions to Perplexity’s knowledge or control.
Several apparently simple questions must be separated:
- Was the content publicly accessible?
- Was it accessible to a particular crawler under Reddit’s rules?
- Was it obtained from Google’s index rather than directly from Reddit?
- Did the method violate copyright law, a contract, an access restriction, or another legal rule?
A “yes” to the first question does not automatically answer the other three. A citation does not prove lawful acquisition, and public visibility does not necessarily grant permission to harvest and commercialize content at scale.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Why the dispute matters beyond Perplexity
Traditional search engines generally send users to source websites. AI answer engines can summarize the source material directly, potentially reducing the need for a user to click through.
Best Value
That creates a difficult economic trade-off. AI systems need high-value human-generated content to answer questions, while publishers depend on visits, advertising, subscriptions, licensing, and attribution. If an AI answer replaces the visit, the publisher may supply the material without receiving equivalent value.
Reddit has also pursued data-licensing revenue, so this is both a legal dispute and a negotiation over who captures the value of user-generated content. Contemporary reporting said Reddit expected more than $200 million over several years from data licensing, but that figure should not be treated as a current financial forecast without later filings or company guidance.
The case also illustrates the expanding role of intermediaries. A website may restrict direct access while search engines, proxy networks, scraping APIs, caches, and data brokers continue to expose pathways to the same information. That makes provenance—where data came from, how it was obtained, and what permissions applied—more important for both AI companies and their customers.
Free tools Windows power users keep installed
One-click scans. No signup required.
What publishers and website owners should take from it
- Do not assume that blocking one crawler blocks every request associated with an AI service.
- Treat declared user agents as identifiers, not proof that all traffic comes from the declared bot.
- Use layered controls such as authentication, rate limits, bot management, firewall rules, monitoring, and access logs where appropriate.
- Remember that preventing direct crawling may not prevent retrieval through search indexes or third-party data providers.
- Document terms, permissions, licenses, and technical restrictions clearly.
- Do not assume that an AI citation means the source received a click, payment, or permission-based access.
Tools such as Cloudflare’s bot-management and web-application-firewall products may help identify or restrict automated traffic, but no single control resolves the legal or economic questions. Technical defenses should be paired with clear licensing policies and monitoring.
So was Perplexity caught “red-handed”?
That phrase accurately reflects Reddit’s characterization of its test, not the legal status of the case.
Reddit produced evidence it says shows that Perplexity’s system reproduced a uniquely marked post that was available to Google but not normally discoverable elsewhere. If proven, that could support the theory that Reddit data reached Perplexity through an indirect scraping route.
But the test does not independently identify the scraper, establish Perplexity’s instructions or knowledge, prove that the content was stored rather than retrieved, or determine whether the conduct violated a specific law. Perplexity denied wrongdoing and disputed Reddit’s framing.
Recommended Free Tools
The most accurate conclusion is therefore narrower than the headline: Reddit says its honeypot exposed indirect access to Reddit content through scraped Google results; Perplexity says the lawsuit mischaracterizes its answer engine and model practices; the technical attribution and legal liability questions remain unresolved.
Quick Recap
Sources
- Reddit’s original complaint, filed October 22, 2025
- Reddit’s first amended complaint, filed February 6, 2026
- Cloudflare’s report on alleged undeclared Perplexity crawlers
- Perplexity’s public response
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

