OpenAI says the company has been using GPT-4 to enforce its content policies and says some of its customers are already using the LLM for content moderation
Can ChatGPT become a content moderator? Semafor https://www.semafor.com/... X: @openai : We've seen great results using GPT-4 for content policy development and content moderation, enabling more consistent labeling, a faster feedback loop for policy refinement, and less involvement from human moderators. Built on top of the GPT-4 API: https://openai.com/... [video] Ethan Mollick / @emollick : Automation traditionally has been about eliminating terrible, repetitive & dangerous jobs. Here is an example of AI doing just that. Social network moderation is a bad job that exposes you to real mental health effects👇 (As a former Wikipedia administrator, I saw it first hand) [image] Rachel Metz / @rachelmetz : content moderation is hard for people (for some reasons obvious such as having to view awful stuff), and it's harder for ai (for various reasons, such as grasping context). so it'll be interesting to see how this goes. from the fabulous @Priyasideas. https://www.bloomberg.com/... Mark Tenenholtz / @marktenenholtz : GPT-4 for content moderation seems impractical at scale given current prices. A more cost-efficient pipeline: • Fine-tune a small classification model (~300M to 7B params, probably) • Set a very high-recall probability threshold • Anything caught by the first stage model... @officiallogank : Big win for a healthy internet: anyone using the @OpenAI API can also use our world class content moderation endpoint to support best in class moderation with minimal setup time. https://openai.com/... Casey Newton / @caseynewton : I talked to some experts about OpenAI's new pitch for platforms to moderate content using GPT-4. They were surprisingly enthusiastic https://www.platformer.news/ ... [image] Sam Houston / @samhouston : @rachelmetz @Priyasideas i once worked as a moderator who had to flag nudity and bullying content. I personally think it will be great if AI can help take over a lot of that work. it feels really demeaning to do that stuff and i still have moments 9-10 years later that are bad memories Andrew Ruiz / @then_there_was : @webdevMason Hopefully one day we'll have an LLM as an intermediary to filter out the truly insane replies. I don't think that future is far away. Could probably be implemented today too with GPT-4 and a custom Twitter client. https://twitter.com/... Reed Albergotti / @reedalbergotti : @cartwheelit @Techmeme you might be able to prompt your chatbot to say something inappropriate. that is what could go wrong. Reed Albergotti / @reedalbergotti : New: OpenAI is now using GPT-4 to moderate ChatGPT. Still not quite as good as a highly-trained team of human moderators, but it's almost there. https://www.semafor.com/... Greg Brockman / @gdb : GPT-4 for content moderation. Very reliable at this use-case (dark blue bars are GPT-4, other bars are well-trained and lightly-trained humans) & speeds up iterating on policies (sometimes literally from months to hours). [image] Forums: r/singularity : Using GPT-4 for content moderation
Context & Ripple Effects
OpenAI is moving moderation from a labor-intensive support function toward an API-enabled model capability. Earlier coverage of its Kenyan contractor arrangement showed the human labeling work behind safety systems; this report positions GPT-4 as a tool to reduce that work in ongoing policy enforcement.
The move also fits OpenAI’s broader pattern of packaging task-specific behavior for users, later reflected in custom GPTs tailored to particular jobs. Moderation is a consequential test case because policy interpretation, not just text classification, is being delegated to the model.
First-order effects
- OpenAI can use GPT-4 to apply its own content rules and offer customers a GPT-4-based moderation workflow, potentially making labeling faster and more consistent across policy changes.
- Customers using the API gain an alternative to moderation processes that rely more heavily on human review; human moderators remain relevant where model judgments need escalation or checking.
Second-order effects
- Platforms and other AI providers face pressure to make safety enforcement configurable, auditable, and quick to update rather than treating moderation as a fixed classifier or outsourced operation.
- The value of human moderation work shifts toward creating policy examples, reviewing difficult cases, and evaluating model decisions—work foreshadowed by the human labeling operation used to improve ChatGPT.
Third-order effects
- If models increasingly enforce policies that humans define, operational AI governance becomes a product capability: organizations will need clear policies and review paths for contested model decisions.
- This points toward a layered safety stack in which specialized models assist reviewers, akin to OpenAI’s later CriticGPT approach to helping human evaluators find errors, rather than a clean replacement of human oversight.
The trend: Content moderation is becoming an LLM-native operational workflow, combining policy authoring, automated first-pass decisions, and targeted human review.