Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Amazon Web Services investigated Perplexity AI in June 2024 after WIRED reported that an AWS-hosted server appeared to scrape publisher websites despite those sites using robots.txt instructions to discourage automated access. The investigation was not a public finding that Perplexity violated AWS rules. Perplexity denied that its own crawler breached those rules, and the public record does not establish a final AWS enforcement decision.

What happened in June 2024?

On June 27, 2024, WIRED reported that Amazon was investigating whether Perplexity used AWS infrastructure to scrape websites that had attempted to block automated crawlers.

WIRED traced an unpublished IP address to an AWS EC2 virtual machine and reported repeated visits from that infrastructure to Condé Nast properties. The reporting also described similar activity involving websites operated by or associated with The Guardian, Forbes and The New York Times.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The concern was not simply that Perplexity used AWS. Cloud providers host many kinds of software and are not automatically responsible for every request made by a customer. The question was whether AWS resources were being used to facilitate activity prohibited by Amazon’s customer-use rules.

What Amazon actually said

Amazon told WIRED that it was investigating information supplied by the publication concerning a possible AWS Terms of Service violation. AWS says its Acceptable Use Policy prohibits illegal or fraudulent activity, violations of other people’s rights, and interference with the security, integrity or availability of computer systems.

The policy also allows AWS to investigate suspected violations and, where appropriate, disable access to resources involved in prohibited activity. That language explains why Amazon could examine conduct occurring on a customer’s EC2 instance. It does not mean AWS had already decided that a violation occurred.

There is no public evidence in the supplied record that AWS suspended or terminated Perplexity’s account, or that it announced a completed finding in the 2024 matter. AWS’s current Service Terms contain investigation and enforcement mechanisms, but current wording should not automatically be treated as identical to the terms in force in June 2024.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What Perplexity said

Perplexity denied that its controlled services violated AWS rules. Its reported position was that PerplexityBot respected robots.txt. The company also distinguished ordinary web crawling from a situation in which a user directly supplied a URL and asked the service to retrieve or summarize it.

Perplexity attributed the unpublished IP address to a third-party crawling or indexing service and did not publicly identify that provider. That claim left control of the relevant software, AWS account and requests as a disputed factual issue.

An IP address can identify infrastructure or an account relationship, but it does not by itself prove which company controlled the crawler or who authorized each request. Establishing that connection would require additional evidence about account ownership, software, logs, contracts or operational control.

Why robots.txt matters—and what it does not prove

robots.txt is a standard way for a website operator to publish instructions for automated crawlers. A site might use a Disallow rule to ask compliant bots not to access a path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It is important to separate five different questions:

  • What instructions did the site publish?
  • Did a crawler choose to follow those instructions?
  • Did the website technically block the request?
  • Did the activity violate an AWS contractual policy?
  • Did the conduct violate copyright, contract, computer-access or other laws?

A crawler ignoring robots.txt is not automatically the same as bypassing a password wall, CAPTCHA, paywall or other technical access control. Nor is robots.txt itself a law, license or security barrier. It can nevertheless be important evidence in a contractual, commercial or legal dispute, especially when combined with the crawler’s behavior, identity and method of access.

Perplexity’s current crawler documentation lists separate PerplexityBot and Perplexity-User agents. The documentation says Perplexity-User may fetch a page in response to a user request and generally ignores robots.txt because the fetch is user-requested.

Perplexity’s current help-center explanation says it will not index full or partial text from sites that disallow it through robots.txt. It also says users previously could ask Perplexity to summarize a blocked URL, but that feature was later disabled, and that agreements with third-party crawlers were updated to require compliance, particularly for news publishers. Those are current company policy statements, not proof of exactly what happened in 2024.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What remains unproven

The public reporting establishes an investigation and a dispute over attribution. It does not establish all of the following:

  • That Perplexity directly operated the AWS-hosted server.
  • That the relevant requests were made by PerplexityBot.
  • That every request ignored a valid robots.txt rule in effect at that time.
  • That the activity bypassed a technical access control.
  • That the conduct was unlawful.
  • That AWS made a final finding that Perplexity violated its rules.
  • That Amazon suspended or banned Perplexity from AWS.

Historical details also matter. A website’s robots.txt file can change, so evaluating a past request requires examining what the file said at the time—not only its current contents. The identity of the user agent, the originating IP address, whether requests were automated or user-triggered, and whether third-party infrastructure was involved are all material facts.

How AWS’s role differs from a publisher’s role

A publisher can object to automated access, restrict traffic technically, pursue a contractual or copyright claim, or argue that an automated system misrepresented its identity. AWS evaluates a separate relationship: whether its customer used AWS services in a way prohibited by AWS’s policies.

That distinction creates a difficult balance for cloud providers. AWS must investigate credible abuse reports and protect the security and availability of systems, while also avoiding the assumption that hosting infrastructure makes it the operator of every application running there. Third-party crawlers, proxy services, browser automation and contractors can further complicate attribution.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

The separate 2025 Comet dispute

The 2024 AWS investigation should not be merged with the later Amazon–Perplexity dispute over Comet.

Date Development Issue
June 2024 AWS investigation reported by WIRED Alleged scraping of publisher websites from AWS-hosted infrastructure.
July 2025 onward Perplexity launched Comet An AI-enabled browser designed to take actions for users, including shopping-related actions.
November 2025 Amazon issued objections and later sued Amazon alleged that Comet agents accessed customer accounts and interacted with Amazon’s store without authorization.

In its public statement and cease-and-desist letter, Amazon alleged that Comet failed to identify itself as an AI agent and disguised automated activity as ordinary browser traffic. The later lawsuit concerned agentic browsing and shopping inside Amazon’s store, not the original AWS-hosted scraping investigation.

Those disputes are related thematically because both involve automated access and the difficulty of identifying AI systems. However, the later allegations do not retroactively prove that Perplexity violated AWS rules in 2024, and the later litigation should not be presented as the outcome of the earlier investigation.

Why the case mattered

The episode highlighted a growing conflict between AI search companies and publishers. AI services need large amounts of web information and may also fetch pages in response to individual users. Publishers, meanwhile, want control over automated access, attribution, licensing and traffic.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It also exposed the limits of voluntary crawler standards. A declared bot identity and robots.txt compliance make activity easier to understand and manage, but generic browser identities, rotating infrastructure and third-party services can make enforcement harder. User-directed fetching adds another complication: a request may be initiated by a user while still being executed automatically by an AI system.

For cloud providers, the issue is part of a broader challenge: determining when infrastructure is being used for ordinary automation and when it is being used for abusive, unlawful or rights-violating conduct. The answer may depend less on the mere fact of scraping than on the target, scale, method, consent signals, technical barriers and control relationships.

Bottom line

AWS investigated allegations that Perplexity-related infrastructure scraped publisher websites despite robots.txt restrictions. Perplexity denied that its controlled crawler violated AWS rules and said the relevant server belonged to a third party. The public record described here does not show a final AWS determination, account suspension or judicial finding that the 2024 activity was unlawful. Amazon’s later lawsuit over Perplexity’s Comet shopping agent was a separate dispute.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.