/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Analysis: Perplexity seems to scrape sites using surreptitious methods, ignoring robots.txt, with a Perplexity-tied machine doing so on Wired and other sites

A WIRED investigation shows that the AI-powered search startup Forbes has accused of stealing its content is surreptitiously scraping—and making things up out of thin air.

Wired

Context & Ripple Effects

The report extends Forbes' earlier complaint about near-verbatim republication and weak attribution from output-level concerns to the methods used to obtain web content. It also arrives as Perplexity was reportedly discussing revenue-sharing arrangements with publishers, making the credibility of its publisher relationship central to the dispute.

Perplexity's subsequent response that it relies on third-party crawlers while not ignoring robots.txt places responsibility for collection practices, and oversight of those vendors, at the center of the issue.

First-order effects

  • Publishers named in the reporting face a more immediate need to assess whether their access controls are being honored and whether Perplexity-generated answers reproduce or misstate their reporting.
  • Perplexity faces intensified scrutiny over both alleged crawl behavior and answer reliability; its third-party crawler explanation becomes a key test of its operational accountability.

Second-order effects

  • Publisher negotiations with AI search providers may shift from commercial terms alone toward auditable crawling practices, attribution, and remedies for inaccurate summaries.
  • Other answer engines and crawler vendors have an incentive to clarify how they interpret site permissions, since publishers may treat opaque collection paths as a reason to restrict access.

Third-order effects

  • If publishers cannot reliably enforce machine-readable access preferences across direct and third-party collection, robots.txt becomes a weaker practical control for the answer-engine market.
  • The dispute points toward AI-search licensing models in which provenance, crawl governance, and output attribution are bundled rather than handled as separate questions.

The trend: AI answer engines are moving from a debate over summarized outputs to a broader contest over who controls web-content access, provenance, and compensation.