In response to plagiarism allegations, Perplexity CEO Aravind Srinivas says the company “is not ignoring” robots.txt, but does rely on third-party web crawlers
* what we do is highly technical, you don't understand — * it wasn't us it was a third party service/contractor/vendor — https://www.fastcompany.com/ ... @bsmall2@mstdn.jp : Automated Plagiarism for BS (not me, the other BS, Harry Frankfurt's) : — > Robots.txt is a single bit of code that's been used since the late 1990s as a way for websites to tell bot crawlers they don't want their data scraped and collected. It was widely accepted as one of the unofficial rules supporting the web... … @johnvoorhees@mastodon.macstories.net : Six days later, what started as a ‘WTF is going on?’ moment for @robb, @viticci, and me has become something else entirely. — The latest from Business Insider, which reports that OpenAI and Anthropic are ignoring publishers' robots.txt files too. — https://www.businessinsider.com/ ... … Bluesky: Jenn Cutter / @jenncutter.bsky.social : Lemme summarize: “We're not ignoring robots.txt! We're just specifically partnering with services that do. No, you can't know their names. It's fine! All you whiners can sit down, shut up, and like it.” [embedded post] X: Dare Obasanjo / @carnage4life : Perplexity's CEO responds to claims they are ignoring robots.txt and crawling websites by saying it's the vendor they use for crawling that's does it not them and robots.txt isn't a law anyway. It's interesting to watch AI companies break the social contracts the web is built on [image] Halvar Flake / @halvarflake : People that ignore robots.txt automatically waive their side of all other agreements, such as ToS or licensing terms. @luke_metro : honestly it's amazing that robots.txt held up as a self-regulation for 25 years [image] LinkedIn: Glenn Gabe : Oh boy, so a mysterious third-party is crawling and indexing your site and feeding that to Perplexity. … Rafael Brown : So Wired has proven that Perplexity ignores protocol and scrapes everything they can find. And Business Insider is saying the same thing about OpenAI and Anthropic. … See also Mediagazer
Context & Ripple Effects
Perplexity’s response follows a report alleging that Perplexity-linked scraping bypassed robots.txt and puts its crawling practices, rather than only the resulting answers, at the center of the dispute. The company’s distinction is that third-party vendors participate in collection and indexing.
The episode matters because robots.txt is a publisher-facing control whose effectiveness depends on every participant in an AI search supply chain honoring it. Later coverage of publishers updating robots.txt to block Anthropic bots shows that crawler identification and compliance were already becoming a contested operating issue.
First-order effects
- Perplexity must account for how its third-party crawling partners collect and index content, since its defense makes vendor behavior directly relevant to the allegations.
- Publishers concerned about unauthorized reuse face a more complex enforcement target: the AI product provider and the third-party crawlers operating on its behalf.
Second-order effects
- AI search companies and crawler vendors face pressure to make bot identities, client relationships, and robots.txt compliance more auditable; publisher blocks are less useful when responsibility is split across intermediaries.
- Publishers may tighten access controls and evaluate terms-of-service enforcement alongside robots.txt, increasing friction between content owners and AI-driven discovery products.
Third-order effects
- The dispute exposes robots.txt as a voluntary, single-layer control that may not map cleanly to multi-vendor AI data pipelines; if such arrangements proliferate, publishers will seek more enforceable technical or commercial permissions.
- AI search is moving toward a governance model in which access to web content is not merely a crawling question but a provenance and accountability question across the full collection chain.
The trend: AI search is forcing a shift from informal crawler etiquette toward verifiable publisher-control and content-provenance systems.