/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Microsoft, OpenAI, Cohere, and others are testing the use of “synthetic data”, as they find generic data from the web is no longer good enough for training LLMs

Microsoft, OpenAI and Cohere experiment with “synthetic data,” as they reach the limits of information created by humans

Financial Times Madhumita Murgia

Context & Ripple Effects

This is an early signal that leading model developers saw generic web corpora as an insufficient foundation for further training gains. Later coverage framed the same constraint as a possible push toward smaller, more specialized models rather than ever-larger general-purpose systems.

Synthetic data is not a universally safe substitute: subsequent researchers warned of model degradation from recursive synthetic training. As models also near the limits of public tests, companies have begun building their own internal benchmarks, making data generation and evaluation increasingly linked.

First-order effects

  • Microsoft, OpenAI, Cohere, and peers must test synthetic-data pipelines alongside web-derived datasets, shifting part of model development from data collection toward controlled data creation and validation.
  • The immediate competitive question becomes whether a lab can produce synthetic examples that improve a target capability without degrading the model's broader behavior.

Second-order effects

  • Demand shifts toward specialized datasets, curation, and evaluation tooling, because synthetic outputs require stronger quality checks than simply adding more web-scale text.
  • Competitors pursuing general-purpose models face a data-supply constraint that can favor teams able to pair generated training material with proprietary tasks, feedback, and internal tests.

Third-order effects

  • If conventional web data remains constrained, the industry may move from a common public-data frontier toward differentiated, task-specific training and evaluation stacks.
  • The key long-run constraint may become data quality and provenance rather than raw volume; the reported collapse risk means synthetic data is more likely to be a managed input than a frictionless replacement for human-created material.

The trend: This is one data point in the shift from web-scale data aggregation to controlled synthetic-data, proprietary-data, and evaluation-driven model development.