Analysis: Perplexity seems to scrape sites using surreptitious methods, ignoring robots.txt, with a Perplexity-tied machine doing so on Wired and other sites
A WIRED investigation shows that the AI-powered search startup Forbes has accused of stealing its content is surreptitiously scraping—and making things up out of thin air.
Context & Ripple Effects
The report extends Forbes' earlier complaint about near-verbatim republication and weak attribution from output-level concerns to the methods used to obtain web content. It also arrives as Perplexity was reportedly discussing revenue-sharing arrangements with publishers, making the credibility of its publisher relationship central to the dispute.
Perplexity's subsequent response that it relies on third-party crawlers while not ignoring robots.txt places responsibility for collection practices, and oversight of those vendors, at the center of the issue.
First-order effects
- Publishers named in the reporting face a more immediate need to assess whether their access controls are being honored and whether Perplexity-generated answers reproduce or misstate their reporting.
- Perplexity faces intensified scrutiny over both alleged crawl behavior and answer reliability; its third-party crawler explanation becomes a key test of its operational accountability.
Second-order effects
- Publisher negotiations with AI search providers may shift from commercial terms alone toward auditable crawling practices, attribution, and remedies for inaccurate summaries.
- Other answer engines and crawler vendors have an incentive to clarify how they interpret site permissions, since publishers may treat opaque collection paths as a reason to restrict access.
Third-order effects
- If publishers cannot reliably enforce machine-readable access preferences across direct and third-party collection, robots.txt becomes a weaker practical control for the answer-engine market.
- The dispute points toward AI-search licensing models in which provenance, crawl governance, and output attribution are bundled rather than handled as separate questions.
The trend: AI answer engines are moving from a debate over summarized outputs to a broader contest over who controls web-content access, provenance, and compensation.