How AI's post-training process suppresses the creativity and whimsicality seen in earlier models, like GPT-2, leading to poor writing from many top AI models
Why can't language models write well? — In a certain, strange way, generative AI peaked with OpenAI's GPT-2 seven years ago.
The AtlanticJasmine Sun
Context & Ripple Effects
The story reframes apparent progress in language models as a post-training trade-off: systems may become more controlled while losing some of the unexpected qualities associated with earlier models. That complicates the earlier critique of GPT-2's superficial and unreliable knowledge—creative-sounding output and dependable output are not the same measure.
The writing question has already produced mixed evidence. A GPT-4 story-idea experiment found gains for some writers but less diversity across the resulting work, while accounts of AI-assisted writing have described generated insights as hollow or approximate.
First-order effects
For people using leading models to draft prose, post-training behavior becomes a direct quality constraint: a model can be capable yet produce writing perceived as flattened or overly conventional.
Model developers targeting writing use cases face a clearer trade-off between alignment toward controlled output and preserving stylistic range or surprise.
Second-order effects
Writing-focused AI products will be pushed to differentiate polished, predictable generation from tools that help users retain voice and generate genuinely varied alternatives.
Evaluations of language models may need to weigh diversity and literary quality alongside reliability, because aggregate capability claims do not settle how useful a model is for creative work.
Third-order effects
If this pattern holds across leading systems, post-training—not just larger base models—will become a central competitive lever in creative AI, fragmenting products by the kinds of expression they preserve or constrain.
The broader market could separate AI used to standardize routine communication from AI designed for exploratory creative work, though the corpus does not establish that this split is inevitable.
The trend:Generative AI is moving from a race over raw model capability toward a contest over how post-training shapes output quality, control, and creative range.
I'd be so interested to know who wrote the “personalized” edits that this writer fed to the chatbot. Because it sure sounds like she input the labor of her past human editors...?
This is a cool example of how you can use AI to help your writing—without relying on it for any actual writing. From @jasminewsun https://www.theatlantic.com/ ... [image]
Always a great read from @jasminewsun. I agree that AI is still a better editor than a writer. But as someone whose most successful post was written by AI, I have to push back on the idea that “LLMs may never be capable of great writing themselves.” I think they already are.
As recently as a few years ago, my impression is, knowledgeable people thought AI would be better at writing than coding. Obviously hasn't turned out that way. I think it all comes down to surprise. Great writing constantly surprises you. Within a single sentence it takes turns
I am generally quite bullish on AI capabilities, so this piece was as much about LLM training as it is about the ineffable, unpredictable qualities that make good writing good [image]
“Modern LLMs are built in a way that is antagonistic to great writing; they are engineered to be rule-following teacher's pets that always have the right answer in hand.” The brilliant @jasminewsun has written a compelling diagnosis of why LLM writing is so tight-lipped:
somehow the same AIs that can do PhD-level math and superhuman coding can only write as well as “a real poet's okay poem” (sama's words, not mine!) I talked to the people training AIs to write about what makes it so hard: new from me for @TheAtlantic: https://www.theatlantic.com/…
“Sam Altman has predicted that large language models will soon be capable of 'fixing the climate, establishing a space colony, and the discovery of all of physics” — but, “might be able to extrude only something equivalent to 'a real poet's okay poem'” @jasmine.bsky.social on wha…
“Chatbots produce meaningless metaphors, endless 'it's not this, but that' constructions, and a cloyingly sycophantic tone” — www.theatlantic.com/technology/ 2...