/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

A look at OpenAI's “red team” of 50 academics and experts, hired in 2022 to look for issues such as toxicity, prejudice, and biases in GPT-4 before its release

Microsoft-backed company asked an eclectic mix of people to ‘adversarially test’ GPT-4, its powerful new language model

Financial Times Madhumita Murgia

Context & Ripple Effects

OpenAI's red team is the company's structural answer to an old criticism: back in 2020, Facebook's AI VP Jerome Pesenti called out GPT-3 for easily outputting toxic language that propagates harmful biases. By 2023, OpenAI had moved from defending its models to hiring 50 outside academics and experts to adversarially attack GPT-4 for toxicity, prejudice, and bias before release.

The red team also sits alongside a parallel external check: weeks earlier, OpenAI let the Alignment Research Center — founded by a former employee — assess GPT-4 for risks including power-seeking behavior. Together they show a lab building pre-release assurance from both hired critics and independent evaluators, a pattern later formalized when OpenAI shipped GPT-Red, an automated red-teaming model for finding prompt injection bugs.

First-order effects

  • The 50 hired academics and experts gain a formal adversarial role against GPT-4 before its release, with their findings on toxicity, prejudice, and biases feeding into OpenAI's launch decisions.
  • OpenAI converts pre-release testing from an internal checkbox into an external audit it can point to, complementing the Alignment Research Center's independent GPT-4 risk assessment.

Second-order effects

  • The GPT-4V paper later published some of the model's residual biases, flaws, and malicious-use cases alongside its safeguards, showing red-team findings becoming part of OpenAI's public documentation rather than private fixes.
  • Rival frontier labs face a rising expectation that major model launches come with disclosed adversarial testing, since OpenAI has normalized both paid expert panels and third-party evaluators as launch hygiene.

Third-order effects

  • If the pattern holds, human red teams become the seed of an institutionalized assurance pipeline — OpenAI's own GPT-Red automates prompt-injection discovery at scale, suggesting the expert-panel model evolves into continuous machine-run testing before deployment.
  • Pre-release adversarial auditing is on track to become a baseline requirement for frontier-model credibility, with labs competing on who tests, how independently, and how much of the findings they publish.

The trend: Frontier-lab safety testing is institutionalizing from ad-hoc human expert panels into scaled, eventually automated, pre-release red-teaming.

Discussion

  • @csetgeorgetown @csetgeorgetown on x
    “The reason why you do operational testing is because things behave differently once they're actually in use in the real environment.” CSET's Heather Frase spoke to @FT's @madhumita29 about red-teaming GPT-4. https://www.ft.com/...
  • @hatr Hakan on x
    “After Andrew White was granted access to GPT-4, the new artificial intelligence system that powers the popular ChatGPT chatbot, he used it to suggest an entirely new nerve agent.” https://www.ft.com/...
  • @madhumita29 Madhumita Murgia on x
    Soon after GPT-4 came out I spoke to @paul_rottger for a story, who sparked the idea to speak to as many “red-teamers” as possible. Here's the fruit of that reporting which hopefully gives some insight into the inner workings of this complex technology - and how it could go wrong…
  • @betweenmyths Ken Craggs on x
    @emollick “White [@andrewwhite01] told the Financial Times he had used GPT-4 to suggest a compound that could act as a chemical weapon...The chatbot then even found a place to make it.” https://www.ft.com/... @FT #AI #GenerativeAI #GPT4 #OpenAI #OpenAIChatGPT #OpenAIRedTeam
  • @agiquality Agi on x
    OpenAI's red team broke ChatGPT, exposing potential dangers of powerful AI systems. Findings led to improved safety measures before the public release of GPT-4. #AI #ChatGPT #OpenAI https://www.ft.com/...
  • @muradahmed Murad Ahmed on x
    Superb reporting - @madhumita29 spoke to OpenAI's “red team” - the experts hired to try and “break” GPT4. Amongst other things, one chemical professor used it to create a nerve agent - in @FT https://www.ft.com/...
  • @financialtimes @financialtimes on x
    We talked to the eclectic mix of people inside OpenAI's ‘red team’ — experts hired to ‘adversarially test’ GPT-4 with probing or dangerous questions https://www.ft.com/...