/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

In response to plagiarism allegations, Perplexity CEO Aravind Srinivas says the company “is not ignoring” robots.txt, but does rely on third-party web crawlers

* what we do is highly technical, you don't understand  —  * it wasn't us it was a third party service/contractor/vendor  —  https://www.fastcompany.com/ ... @bsmall2@mstdn.jp : Automated Plagiarism for BS (not me, the other BS, Harry Frankfurt's) :  —  > Robots.txt is a single bit of code that's been used since the late 1990s as a way for websites to tell bot crawlers they don't want their data scraped and collected.  It was widely accepted as one of the unofficial rules supporting the web... … @johnvoorhees@mastodon.macstories.net : Six days later, what started as a ‘WTF is going on?’ moment for @robb, @viticci, and me has become something else entirely.  —  The latest from Business Insider, which reports that OpenAI and Anthropic are ignoring publishers' robots.txt files too.  —  https://www.businessinsider.com/ ... … Bluesky: Jenn Cutter / @jenncutter.bsky.social : Lemme summarize: “We're not ignoring robots.txt!  We're just specifically partnering with services that do.  No, you can't know their names.  It's fine!  All you whiners can sit down, shut up, and like it.”  [embedded post] X: Dare Obasanjo / @carnage4life : Perplexity's CEO responds to claims they are ignoring robots.txt and crawling websites by saying it's the vendor they use for crawling that's does it not them and robots.txt isn't a law anyway. It's interesting to watch AI companies break the social contracts the web is built on [image] Halvar Flake / @halvarflake : People that ignore robots.txt automatically waive their side of all other agreements, such as ToS or licensing terms. @luke_metro : honestly it's amazing that robots.txt held up as a self-regulation for 25 years [image] LinkedIn: Glenn Gabe : Oh boy, so a mysterious third-party is crawling and indexing your site and feeding that to Perplexity. … Rafael Brown : So Wired has proven that Perplexity ignores protocol and scrapes everything they can find.  And Business Insider is saying the same thing about OpenAI and Anthropic. … See also Mediagazer

Fast Company Mark Sullivan

Context & Ripple Effects

Perplexity’s response follows a report alleging that Perplexity-linked scraping bypassed robots.txt and puts its crawling practices, rather than only the resulting answers, at the center of the dispute. The company’s distinction is that third-party vendors participate in collection and indexing.

The episode matters because robots.txt is a publisher-facing control whose effectiveness depends on every participant in an AI search supply chain honoring it. Later coverage of publishers updating robots.txt to block Anthropic bots shows that crawler identification and compliance were already becoming a contested operating issue.

First-order effects

  • Perplexity must account for how its third-party crawling partners collect and index content, since its defense makes vendor behavior directly relevant to the allegations.
  • Publishers concerned about unauthorized reuse face a more complex enforcement target: the AI product provider and the third-party crawlers operating on its behalf.

Second-order effects

  • AI search companies and crawler vendors face pressure to make bot identities, client relationships, and robots.txt compliance more auditable; publisher blocks are less useful when responsibility is split across intermediaries.
  • Publishers may tighten access controls and evaluate terms-of-service enforcement alongside robots.txt, increasing friction between content owners and AI-driven discovery products.

Third-order effects

  • The dispute exposes robots.txt as a voluntary, single-layer control that may not map cleanly to multi-vendor AI data pipelines; if such arrangements proliferate, publishers will seek more enforceable technical or commercial permissions.
  • AI search is moving toward a governance model in which access to web content is not merely a crawling question but a provenance and accountability question across the full collection chain.

The trend: AI search is forcing a shift from informal crawler etiquette toward verifiable publisher-control and content-provenance systems.

Discussion

  • @camwilsonreporter Cameron Wilson on threads
    CEO of Perplexity, the AI search engine company caught scraping web content even after its creator opted-out, says opt-out framework (robots.txt) isn't “legally binding” It's like when you say something shitty, have no excuse & fall back on saying “it's freedom of speech” https:/…
  • @glenngabe Glenn Gabe on threads
    Oh boy, so a mysterious third-party is crawling and indexing your site and feeding that to Perplexity.  How wonderful :) -> In response to plagiarism allegations, Perplexity CEO Aravind Srinivas says the company “is not ignoring” robots.txt, but does rely on third-party web crawl…
  • @johnvoorhees@mastodon.macstories.net @johnvoorhees@mastodon.macstories.net on mastodon
    Six days later, what started as a ‘WTF is going on?’ moment for @robb, @viticci, and me has become something else entirely.  —  The latest from Business Insider, which reports that OpenAI and Anthropic are ignoring publishers' robots.txt files too.  —  https://www.businessinsider…
  • @jenncutter.bsky.social Jenn Cutter on bluesky
    Lemme summarize: “We're not ignoring robots.txt!  We're just specifically partnering with services that do.  No, you can't know their names.  It's fine!  All you whiners can sit down, shut up, and like it.”  [embedded post]
  • @halvarflake Halvar Flake on x
    People that ignore robots.txt automatically waive their side of all other agreements, such as ToS or licensing terms.
  • @luke_metro @luke_metro on x
    honestly it's amazing that robots.txt held up as a self-regulation for 25 years [image]