Amazon is investigating Perplexity over whether the AI search startup is violating AWS rules by scraping websites that attempted to prevent it from doing so
AWS hosted a server linked to the Bezos family- and Nvidia-backed search startup that appears to have been used to scrape the sites …
Wired
Context & Ripple Effects
The AWS inquiry follows reporting that a Perplexity-linked machine appeared to bypass website preferences, including robots.txt, in its crawling activity reported use of surreptitious scraping methods. Perplexity had simultaneously positioned publisher revenue-sharing as part of its response to concerns over AI search use of news content publisher revenue-sharing discussions.
The issue matters because AWS is both the startup’s infrastructure provider and the party whose service rules are implicated. That makes the dispute about more than search-result attribution: it tests whether hosting platforms will police how AI customers obtain web data.
First-order effects
Perplexity faces scrutiny over whether an AWS-hosted server was used to access sites that had sought to block scraping, putting its crawler operations and AWS-account compliance under review.
Website operators that rely on robots.txt or related blocking measures gain a prominent test of whether those signals can trigger action through a crawler’s cloud provider. Perplexity had said it relies partly on third-party crawlers in explaining its crawler practices.
Second-order effects
AI search providers and their crawler vendors may need clearer records of crawler identity, site permissions, and responsibility when infrastructure accounts are used for collection.
Publishers have more reason to pair technical controls with direct commercial or enforcement channels, rather than treating robots.txt alone as a sufficient safeguard.
Third-order effects
If cloud providers increasingly treat unauthorized collection as an account-compliance issue, enforcement can shift from an uneven publisher-by-publisher contest to a chokepoint in AI-search infrastructure.
The lasting question is whether technical opt-outs, contracts, and platform rules converge into a workable permission framework; this inquiry alone does not establish that outcome.
The trend: AI search is moving toward a more enforceable web-data permission regime, with publishers, crawlers, and cloud hosts each becoming part of the compliance chain.
Regarding my last retweet, I did dabble with the likes of Co-pilot when researching articles I've been writing, but always felt uncomfortable doing so. The article in question (below) gives me even more reason not to use it. Funny how these tech bros think they're above copyrig…
»i'm not the one who's telling you the foundation of perplexity is lying to dodge established principles that hold up the web. its ceo is. that's clarifying about the actual value proposition of ›answer engines.‹ perplexity cannot generate actual information on its own and reli…
Amazon's cloud division has launched an investigation into Perplexity AI, to determine whether the AI search startup is violating AWS rules by scraping websites that tried to prevent it from doing so https://www.wired.com/... via @AndrewCouts @dmehro
NEW from @AndrewCouts and @dmehro: Amazon Web Services is investigating Perplexity following a WIRED investigation into the “answer engine.” Clients must, Amazon tells WIRED, respect robots.txt: https://www.wired.com/...
Perplexity is basically saying they follow the robots.txt protocol in relation to “normal” web crawling, but they make an exception when there's a direct user request to access a URL. Is there anything to really do here, especially if no technical enforcement for the publisher?
NEW: Perplexity was using a secret AWS-hosted virtual machine to scrape the web. Major newsroom we spoke to observed its IP address in their server logs. Now Amazon is investigating. w/ @AndrewCouts https://www.wired.com/...
“The machine associated with Perplexity appears to be engaged in widespread crawling of news websites that forbid bots from accessing its content. Spokespeople for the Guardian, Forbes, and The New York Times also say they detected the IP address on its servers multiple times.”