Sources: OpenAI developed a watermarking method for detecting text written by ChatGPT with 99.9% reliability, but its launch has been mired in internal debates
Technology that can detect text written by artificial intelligence with 99.9% certainty has been debated internally for two years
Wall Street Journal
Context & Ripple Effects
OpenAI's reported method follows an earlier period in which it was exploring statistical watermarking for AI-generated text while the technical limits of such systems were under scrutiny. The company also previously retired a separate AI-text classifier because of accuracy problems, shifting attention toward more effective provenance techniques.
The significance is less the reported detection performance than the unresolved decision to ship it. A watermark built into ChatGPT's output would be a different accountability mechanism from a post-hoc classifier, but internal debate has kept that mechanism from becoming available.
First-order effects
OpenAI must continue to weigh the product and policy costs of release against the value of giving users and institutions a native way to identify ChatGPT-originated text.
Until a launch decision is made, ChatGPT users and downstream readers do not have an OpenAI-provided watermark signal to use in their own review workflows.
Second-order effects
Schools, publishers, and employers seeking evidence of AI authorship remain dependent on their own policies or third-party detection approaches rather than a platform-level provenance signal.
A release could make provenance a competitive feature for generative-AI products; continued delay leaves competitors room to differentiate on disclosure and traceability.
Third-order effects
The episode points to a broader split between detecting AI text after the fact and designing provenance into generation itself; the latter is likely to become more important if providers can deploy it without undermining product adoption.
Whether watermarking becomes a practical industry standard will depend on adoption across tools and on how reliably signals survive real-world editing, not solely on a provider's reported performance.
The trend: Generative-AI providers are moving from unreliable after-the-fact detection toward embedded provenance systems, while debating the commercial and governance trade-offs of making those systems visible.
OpenAI faces the classic big tech tension of trust & safety versus growth. On one hand, you can create a tool that detects when people use ChatGPT to cheat on their homework. But then everyone who wants to cheat on their homework will NOT use ChatGPT. User growth goes 📉
There's so much more in this story by me & Matt Barnum about their decision-making process & the tension between their stated goals of transparency + helping educators and their need to grow & destigmatize AI's use. Please read and let me know what you think.
One reason: an internal April 2023 survey showing nearly a third of ChatGPT users would use it less if it deployed watermarking. This survey loomed large even after internal tests in 2024 showed watermarking didn't hurt ChatGPT's output. Here's how OpenAI itself described the s…
OpenAI says they don't want to ostracise their users with watermarks to “protect them”. But I got confirmation that they send Oxford the complete listing of the staff and students using its services (Oxford argued it's not a breach of GDPR).
If this is true it's remarkably irresponsible (hi @responsibleaiuk) Being able to detect AI text would be staggeringly useful in education. as well as to filter AI spam clickbait etc https://www.theverge.com/...
This caught my eye: “That same month, OpenAI surveyed ChatGPT users and found 69% believe cheating detection technology would lead to false accusations of using AI. Nearly 30% said they would use ChatGPT less if it deployed watermarks and a rival didn't.”
OpenAI has a tool that detects ChatGPT written content reliably and could solve numerous problems, but won't release it because it would affect their bottom line. They're lying to you when they say they care about helping humanity, and this is proof. https://www.wsj.com/... [imag…
So basically ChatGPT has for over a year been sitting on a software that can effectively test for cheating, but hasn't released it largely because without cheaters their business would suffer. https://www.wsj.com/...
Now that multiple vendors offer highly capable LLMs, it seems to me that watermarking is pretty much a dead-end - if someone wants to cheat they have multiple options for LLMs that don't watermark, which means vendors have little incentive to add watermarks to their own products