/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Researchers: DeepSeek's R1 failed to detect or block any of 50 randomly selected malicious prompts; Adversa says DeepSeek's restrictions can easily be bypassed

Unit 42 researchers recently revealed two novel and effective jailbreaking … Victor Tangermann / Futurism : DeepSeek Failed Every Single Security Test, Researchers Found Ivan Novikov / Wallarm : Analyzing DeepSeek's System Prompt: Jailbreaking Generative AI Erin Swanson / Enkrypt AI : AI Race Between U.S. and China Takes a Dark Turn as Red Teaming Report Uncovers Critical Safety Failures Akshaya Asokan / PaymentSecurity.io : DeepSeek AI Models Vulnerable to JailBreaking Markus Kasanmascheff / WinBuzzer : DeepSeek's AI Security Under Fire: 100% Jailbreak Success Exposes Critical Flaws Radhika Rajkumar / ZDNET : Deepseek's AI model proves easy to jailbreak - and worse Zeyi Yang / Wired : Here's How DeepSeek Censorship Actually Works—and How to Get Around It Tushar Subhra Dutta / Cyber Security News : New Jailbreak Techniques Expose DeepSeek LLM Vulnerabilities, Enabling Malicious Exploits Bluesky: Corey Quinn / @quinnypig.com : Feature, not bug.  I've had about enough of tools trying to protect me from myself.  Let me use the chainsaw to lop off a leg, please.  [embedded post] @seaks : That doesn't seem ideal [embedded post] Andrew Couts / @couts : NEW: Cisco and UPenn researchers tested 50 well-known jailbreaks against DeepSeek's AI chatbot, including those related to misinformation, cybercrimes, and other illegal activity.  It stoped exactly zero of them. @mattburgess1.bsky.social and @lhn.bsky.social report: www.wired.com/story/deepse... Mastodon: Michael Veale / @mikarv@someone.elses.computer : Another article which does not properly distinguish between open source models (guardrails scientifically very hard to make robust) and the API as a service (model is in a moderation software stack).  Notes that Llama 3.1 failed in almost the same way as DeepSeek.  What is the actual threat vector? https://www.wired.com/... LinkedIn: Eric Wenger : Eye-opening findings from Cisco and University of Pennsylvania re security and safety of DeepSeek AI model after 100% of malicious prompts randomly selected … Sam Rubin : New DeepSeek research just published from Palo Alto Networks Unit 42 shows the model is vulnerable to jailbreaking, allowing it to generate harmful content with minimal effort or expertise. … Forums: r/technews : DeepSeek's Safety Guardrails Failed Every Test Researchers Threw at Its AI Chatbot

Wired

Context & Ripple Effects

DeepSeek's rapid rise was framed in related coverage as an open-research challenger willing to share breakthroughs, while its chatbot had already drawn scrutiny for an 83% failure rate on news-related reliability tests. This report shifts the concern from answer quality to whether model-level safeguards hold under adversarial use.

The findings also land amid evidence that jailbreaking is not unique to one vendor: Best-of-N jailbreaking research showed black-box attacks can defeat safeguards across frontier systems and modalities. DeepSeek is therefore a salient case of a broader deployment-security problem, not proof of a unique technical category.

First-order effects

  • DeepSeek R1 users and API integrators cannot treat the model's stated restrictions as a dependable guardrail against malicious requests; they need to apply their own access controls, monitoring, and output checks.
  • DeepSeek faces immediate pressure to strengthen refusal behavior and test it against adversarial prompting, while Adversa and Unit 42's findings give enterprise evaluators concrete grounds to scrutinize the model before deployment.

Second-order effects

  • Enterprise buyers may make safety evaluations and red-team results a more explicit procurement criterion, raising the cost of adopting fast-moving models whose capability claims are easier to assess than their misuse controls.
  • Rival model providers and security vendors gain an incentive to publish comparable jailbreak-resistance evidence, as a single model's failures make safety assurance a competitive and operational differentiator.

Third-order effects

  • If repeatable jailbreaks remain common, safety will increasingly be treated as a system-design issue—combining model behavior with adversarial testing methods, application controls, and governed access—rather than a promise embedded in a model's policy layer.
  • The episode strengthens the case for frontier-model access governance in which organizations evaluate the enforcement surrounding a model, not only its benchmark performance; the degree of formal oversight remains uncertain.

The trend: Generative-AI competition is moving from headline model capability toward proving that deployment controls remain effective against routine adversarial prompting.

Discussion

  • @quinnypig.com Corey Quinn on bluesky
    Feature, not bug.  I've had about enough of tools trying to protect me from myself.  Let me use the chainsaw to lop off a leg, please.  [embedded post]
  • @seaks @seaks on bluesky
    That doesn't seem ideal [embedded post]
  • @couts Andrew Couts on bluesky
    NEW: Cisco and UPenn researchers tested 50 well-known jailbreaks against DeepSeek's AI chatbot, including those related to misinformation, cybercrimes, and other illegal activity.  It stoped exactly zero of them. @mattburgess1.bsky.social and @lhn.bsky.social report: www.wired.co…
  • r/technews r on reddit
    DeepSeek's Safety Guardrails Failed Every Test Researchers Threw at Its AI Chatbot