Microsoft, OpenAI, Cohere, and others are testing the use of “synthetic data”, as they find generic data from the web is no longer good enough for training LLMs
Microsoft, OpenAI and Cohere experiment with “synthetic data,” as they reach the limits of information created by humans
Context & Ripple Effects
This is an early signal that leading model developers saw generic web corpora as an insufficient foundation for further training gains. Later coverage framed the same constraint as a possible push toward smaller, more specialized models rather than ever-larger general-purpose systems.
Synthetic data is not a universally safe substitute: subsequent researchers warned of model degradation from recursive synthetic training. As models also near the limits of public tests, companies have begun building their own internal benchmarks, making data generation and evaluation increasingly linked.
First-order effects
- Microsoft, OpenAI, Cohere, and peers must test synthetic-data pipelines alongside web-derived datasets, shifting part of model development from data collection toward controlled data creation and validation.
- The immediate competitive question becomes whether a lab can produce synthetic examples that improve a target capability without degrading the model's broader behavior.
Second-order effects
- Demand shifts toward specialized datasets, curation, and evaluation tooling, because synthetic outputs require stronger quality checks than simply adding more web-scale text.
- Competitors pursuing general-purpose models face a data-supply constraint that can favor teams able to pair generated training material with proprietary tasks, feedback, and internal tests.
Third-order effects
- If conventional web data remains constrained, the industry may move from a common public-data frontier toward differentiated, task-specific training and evaluation stacks.
- The key long-run constraint may become data quality and provenance rather than raw volume; the reported collapse risk means synthetic data is more likely to be a managed input than a frictionless replacement for human-created material.
The trend: This is one data point in the shift from web-scale data aggregation to controlled synthetic-data, proprietary-data, and evaluation-driven model development.