/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A study finds that assigning ChatGPT a persona using its API, like “a bad person” or a certain historical figure, increases the chatbot's toxicity sixfold

It's no secret that OpenAI's viral AI-powered chatbot, ChatGPT, can be prompted to say sexist, racist and pretty vile things.

TechCrunch Kyle Wiggers

Context & Ripple Effects

The study lands in a crowded lane of ChatGPT safety research. Two weeks earlier, Age of AI's filter-free FreedomGPT showed what a chatbot looks like with no guardrails at all, and months later researchers demonstrated that a long character suffix could bypass the guardrails on ChatGPT, Bard, and Claude alike (suffix-attack bypasses). What's different here is the vector: not a clever jailbreak but OpenAI's own API feature for assigning personas, meaning the toxicity comes from a supported configuration, not an attack.

First-order effects

  • Developers building products on the API — custom assistants, historical-figure bots, character apps — inherit a sixfold toxicity spike whenever their persona instructions push the model toward hostile behavior, and OpenAI's usage policies put the moderation burden on them.

Second-order effects

Third-order effects

  • If persona-driven toxicity proves systematic across providers, safety evaluation shifts from testing individual prompts to testing deployment configurations, pushing toward formal governance layers over how AI systems are configured and distributed rather than just what users type.

The trend: AI safety is migrating upstream from blocking bad prompts to governing how models are configured and deployed — personas, GPTs, and API defaults are becoming the new attack surface.

Discussion

  • @martinsfp @martinsfp on x
    Well that's nice: Researchers discover a way to make ChatGPT consistently toxic https://artifact.news/...
  • @neuroecology Adam J Calhoun on x
    ngl, spending all your time trying to make chatbots say horrible things sounds like kind of a fun job I still feel like they could have waited a week and just saw what the chan's were up to though https://techcrunch.com/... https://twitter.com/...
  • @jordannovet Jordan Novet on x
    ‘Furthermore, we find concerning patterns where specific entities (e.g., certain races) are targeted more than others (3× more) irrespective of the assigned persona, that reflect inherent discriminatory biases in the model.’ https://arxiv.org/...
  • @ameetdeshpande_ Ameet Deshpande on x
    Large language models and chatbots are being ubiquitously used. But are they safe? In our large-scale toxicity analysis of ChatGPT, we find that assigning it a “persona” significantly increases its toxicity (up to 6X). Paper: https://www.shorturl.at/anpzY Blog: https://www.shortu…
  • @danieljkelley Daniel Kelley on x
    Oh dear. “Depending on the persona assigned to ChatGPT, its toxicity can increase up to 6x, with outputs engaging in incorrect stereotypes, harmful dialogue, and hurtful opinions.” https://arxiv.org/...