A behind-the-scenes look at how OpenAI's GPT-2 predictive text algorithm works, which can be “fine-tuned” to write phony customer reviews or even news articles
\u003clink href="https://www.newyorker.com/ projects/interactive/2019/191014-seabrook/css/ header_override.css" …
Context & Ripple Effects
This 2019 New Yorker explainer lands mid-arc in OpenAI's text-generation story: GPT-2's ability to be fine-tuned into a phony-review or fake-news engine is exactly the capability that critics at The Gradient soon argued rests on superficial, unreliable knowledge — fluent output without grounded understanding.
The piece also reads as an early warning for what followed: GPT-3 reframed the same architecture as a step toward language understanding, and by the time OpenAI opened consumer creation tools, the abuse pattern this article demonstrated had reappeared at platform scale in the GPT Store's impersonation, copyright-infringing, and jailbreaking GPTs.
First-order effects
- Review platforms and news publishers are directly exposed: anyone with the model and fine-tuning instructions can mass-produce customer reviews and articles indistinguishable in style from human writing.
Second-order effects
- As generation gets easier, verification becomes the scarce input — pushing platforms toward provenance and authenticity tooling, and giving critics of these models' reliability (the superficial-knowledge critique) more evidence when outputs mislead.
- OpenAI's own product decisions absorb the lesson: the guardrail questions around its later consumer surfaces, like the custom GPT builder Altman wants to keep simple, are the same open-access-versus-abuse tradeoff this article surfaced with fine-tuning.
Third-order effects
- If each capability jump — GPT-2 to GPT-3 to GPT-4 — widens both legitimate use and misuse, the industry's structural response is trust infrastructure: provenance standards, detection, and platform moderation becoming as central to text AI as the models themselves.
The trend: Text generation is scaling from research demo to consumer platform faster than verification mechanisms, making synthetic-content abuse a recurring cost of every OpenAI capability release.