Researchers: the guardrails on ChatGPT, Bard, and Claude can be bypassed by adding a long suffix of characters to prompts, generating false and toxic responses
A new report indicates that the guardrails for widely used chatbots can be thwarted, leading to an increasingly unpredictable environment for the technology.
Context & Ripple Effects
This finding lands on top of an already shaky trust story. The same chatbots were flagged in late 2022 for a hallucination problem — reshaping what they learned without regard for truth — and by October 2023 research showed ChatGPT, GPT-4, Bard, and Claude producing harmful, race-based answers to basic medical questions. What the new report adds is different in kind: the failures are not incidental model behavior but a demonstrated bypass of the guardrails themselves, via a long suffix of characters appended to ordinary prompts.
That distinction matters because it converts 'the models are sometimes wrong' into 'the controls meant to stop harm can be deliberately defeated' — across OpenAI's, Google's, and Anthropic's systems at once. The later record bears this out as a durable pattern rather than a one-off: by late 2025 researchers were still finding ways around chatbot safety features, this time with prompts framed as riddle-like poems.
First-order effects
- OpenAI, Google, and Anthropic each face the immediate task of patching a class of attack that works identically across all three chatbots, meaning none of them can treat this as a competitor-specific flaw.
- Enterprises and developers building on these APIs inherit the exposure directly: any user-supplied prompt can carry the suffix, so toxic or false output can surface in production products without the deployer doing anything wrong.
Second-order effects
- Safety claims become a competitive battleground — whichever lab demonstrates faster patching and independent red-team audits gains an enterprise-sales edge over rivals whose guardrails just failed the same test.
- Buyers of chatbot-powered services start demanding contractual assurances and monitoring layers, shifting spend toward vendors and tooling that can detect adversarial inputs rather than trusting the model's built-in filters.
Third-order effects
- If every published bypass (suffixes now, poetic framings later) is followed only by incremental patches, prompt-level guardrails get exposed as structurally insufficient, strengthening the case for external auditing requirements and regulation rather than vendor self-certification.
- The recurring cycle — capability launch, harm documented, bypass published, patch shipped — points toward AI reliability being treated like security: a permanent adversarial discipline with disclosure norms, not a feature that ships once.
The trend: AI safety is settling into a continuous adversarial arms race in which static guardrails on consumer chatbots are repeatedly defeated, pushing accountability from vendor promises toward audited, regulated practice.