OpenAI says GPT-4 poses “at most” a slight risk of helping people create biological threats, per the company's early tests to evaluate “catastrophic” LLM risks
Mark Zuckerberg; Struggling Startups Are Looking For the Exits Michael Nuñez / VentureBeat : OpenAI study reveals surprising role of AI in future biological threat creation Tom Carter / Business Insider : ChatGPT probably won't help create biological weapons, OpenAI says Vish Gain / Silicon Republic : GPT-4 ‘mildy useful’ in creating bioweapons, says ChatGPT X: Tolga Bilge / @tolgabilge_ : It's good to see that this is something that is being worked on. I am unsure to what extent I agree with design principle 3: “The risk from AI should be measured in terms of improvement over existing resources.” In the specific case of open-source, where models can be run on an... Trevor Blackwell / @tlbtlbtlb : I'm glad this experiment was done. Seems like a good test for new models before releasing them. Steven Adler / @sjgadler : I'm proud of the investments we've made here: Developing a careful, rigorous protocol that will continue to be useful into the future. Nathan Benaich / @nathanbenaich : or rather, the result is not statistically significant and can be due to noise. doesn't say what these error bars represent either Aleksander Madry / @aleks_madry : People worry about AI boosting biological threat creation, but how would we know how real this risk is? Here is what we have done in this context so far: Jack / @jack24dd30 : tyler cowen said that he's far more worried about LLMs simply helping terrorist groups run more efficiently and have better organization lol Tejal Patwardhan / @tejalpatwardhan : latest from preparedness @ openai: gpt4 at most mildly helps with biothreat creation. method: get bio PhDs in a secure monitored facility. half try biothreat creation w/ (experimental) unsafe gpt4. other half can only use the internet. so far, gpt4 ≈ internet... but we'll... Greg Brockman / @gdb : Evaluations for LLM-assisted biological threat creation. Current models not very capable at this task, but we want to be ahead of the curve for assessing this and other potential future risk areas: @openai : We are building an early warning system for LLMs being capable of assisting in biological threat creation. Current models turn out to be, at most, mildly useful for this kind of misuse, and we will continue evolving our evaluation blueprint for the future. https://openai.com/... LinkedIn: Aleksander Madry : As part of our Preparedness effort, we are sharing some of our early work assessing LLMs and biological threat creation risk. … Forums: Hacker News : Building an early warning system for LLM-aided biological threat creation r/technology : Mistral CEO confirms ‘leak’ of new open source AI model nearing GPT-4 performance
Context & Ripple Effects
OpenAI’s controlled comparison of experimental GPT-4 with ordinary internet access establishes an early baseline: the model was only mildly useful to bio PhDs, while the company began building an early-warning and evaluation framework for capability changes. That matters because broader GPT-4 coverage had already emphasized unusually wide task performance, including claims of near-human results across professional tasks.
The later arc makes this baseline more consequential rather than dispositive. OpenAI subsequently warned that upcoming systems could raise bioweapon-assistance risk, while Microsoft researchers reported models could help design agents that evade DNA-order screening; these developments put pressure on the limits of static, one-time evaluations.
First-order effects
- OpenAI can point to controlled testing indicating that GPT-4 did not materially outperform existing online information for this task, supporting its assessment that current-model risk is limited.
- The company must maintain and refine its early-warning system and evaluation blueprint, since the reported result is explicitly a capability-specific assessment rather than a permanent safety conclusion.
Second-order effects
- Frontier-model developers face a clearer expectation to test whether new capabilities add meaningful harmful assistance beyond what users can already obtain online, not merely to demonstrate generic model safety.
- Biosecurity evaluators and model providers gain a shared benchmark question—incremental uplift over internet access—but criticism over statistical significance means methodology will be as important as the headline risk rating.
Third-order effects
- If model capability continues to improve, bio-risk governance is likely to shift from broad assurances toward recurring, model-specific evaluations tied to deployment and access decisions.
- The emerging challenge is operational: evaluations must detect when assistance becomes materially more useful, especially as research later identified potential evasion of DNA-screening systems, without treating early low-risk results as a durable guarantee.
The trend: This is an early data point in the move toward continuous dual-use capability evaluations as frontier AI models become more capable and widely accessible.