/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Facebook, Instagram, Craigslist, Tumblr, the NYT, the FT, The Atlantic, Vox Media, USA Today, Condé Nast, and more block Apple's Applebot-Extended AI crawler

This summer, Apple gave websites more control over whether the company could train its AI models on their data.

Wired Kate Knibbs

Context & Ripple Effects

The move extends a publisher-led access-control pattern: major news organizations had already blocked OpenAI’s GPTBot, and a January survey found that most leading US news outlets were blocking AI crawlers. Apple is now encountering the same boundary around web content used for model development.

It matters because the blocked set spans publishers, social platforms, and classifieds, rather than a single content category. That makes crawler permission an operational constraint for Apple’s AI effort, not merely a dispute with one news outlet.

First-order effects

  • Applebot-Extended cannot collect material from the participating sites for the AI-training purpose those sites have opted out of, reducing Apple’s directly accessible web corpus.
  • The blocking organizations retain control over whether their content is available to Apple’s AI systems, including the New York Times, Condé Nast, Facebook, and Instagram.

Second-order effects

  • Other sites that already restrict AI crawlers have a clearer precedent to apply the same controls to Apple, making crawler permissions a more uniform publisher policy rather than a provider-specific exception.
  • AI developers seeking broad, current web coverage must work around a patchwork of exclusions, increasing the value of content they can obtain through permitted sources or direct arrangements.

Third-order effects

  • If large platforms and publishers continue to opt out, training-data access will become more fragmented across AI providers; model quality and coverage may increasingly reflect each company’s distribution and content-access position.
  • Robots-based crawler controls are becoming a practical layer of negotiation over content’s use as model input, though their long-term force will depend on crawler compliance and whether more formal commercial or regulatory frameworks emerge.

The trend: AI training is shifting from open-web collection toward publisher-controlled, increasingly fragmented access to content.

Discussion

  • @knibbs @knibbs on x
    Apple's new AI agent has been out for just a few months, but major publishers (NYT, FT, Atlantic, Conde Nast) and platforms (FB, Insta, Tumblr) are already blocking it https://www.wired.com/...
  • @palewire Ben Welsh on x
    You can find me quoted in @Knibbs' @Wired story on the publishers moving to block Apple's new AI scraper. https://www.wired.com/... [image]