/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

In response to plagiarism allegations, Perplexity CEO Aravind Srinivas says the company “is not ignoring” robots.txt, but does rely on third-party web crawlers

The AI search startup Perplexity is in hot water in the wake of a Wired investigation revealing that the startup …

Fast Company Mark Sullivan

Context & Ripple Effects

The dispute escalated from Forbes' report that Perplexity Pages reproduced story passages with inadequate attribution to a stated effort to pursue revenue-sharing arrangements with publishers. Wired then alleged that a Perplexity-linked machine bypassed robots.txt on Wired and other sites.

Srinivas' response narrows the central question: not merely whether Perplexity honors publisher instructions, but whether responsibility for collection practices extends to the third-party crawlers supplying its service. That distinction matters because the earlier Wired scrutiny of alleged robots.txt circumvention put the technical collection chain at the center of the controversy.

First-order effects

  • Perplexity must defend both its own crawling behavior and the compliance practices of third-party crawlers it relies on, rather than treating robots.txt compliance as a purely internal issue.
  • Publishers challenging Perplexity's use of their work gain a more specific point of inquiry: the provenance and permissions of content obtained through external collection partners.

Second-order effects

  • Any publisher revenue-sharing discussions face a harder negotiation over enforceable controls, attribution, and accountability for intermediaries—not simply compensation for use.
  • AI search rivals that depend on external web-data suppliers may face similar pressure to document crawler behavior and clarify who bears responsibility when publisher restrictions are disputed.

Third-order effects

  • The episode points toward publisher controls becoming a supply-chain governance issue for AI search: product companies may need to ensure that contractors and data partners honor the same access rules they publicly endorse.
  • If publisher objections continue to focus on both copying and crawler conduct, commercial licensing and technical access controls are likely to become more tightly linked, though the extent will depend on publishers' leverage and platform responses.

The trend: AI search is moving from a debate over output attribution toward scrutiny of the full content-acquisition chain, including the intermediaries that gather source material.

Discussion

  • @johnvoorhees@mastodon.macstories.net @johnvoorhees@mastodon.macstories.net on mastodon
    Six days later, what started as a ‘WTF is going on?’ moment for @robb, @viticci, and me has become something else entirely.  —  The latest from Business Insider, which reports that OpenAI and Anthropic are ignoring publishers' robots.txt files too.  —  https://www.businessinsider…
  • @jenncutter.bsky.social Jenn Cutter on bluesky
    Lemme summarize: “We're not ignoring robots.txt!  We're just specifically partnering with services that do.  No, you can't know their names.  It's fine!  All you whiners can sit down, shut up, and like it.”  [embedded post]
  • @luke_metro @luke_metro on x
    honestly it's amazing that robots.txt held up as a self-regulation for 25 years [image]