Google DeepMind releases its Frontier Safety Framework, a set of protocols for analyzing and mitigating future risks posed by advanced AI models
The Scoop — Preparing for a time when artificial intelligence is so powerful that it can pose a serious, immediate threat to people …
Semafor Reed Albergotti
Related Coverage
- Google Introduces Frontier Safety Framework to Identify and Mitigate Future AI Risks Maginative · Chris McKay
- Introducing the Frontier Safety Framework Google DeepMind
Discussion
-
@shanelegg
Shane Legg
on x
Today our team @GoogleDeepMind announced the Frontier Safety Framework — a set of exploratory protocols to proactively identify future AI capabilities, and put in place mechanisms to detect them. Great efforts from @ancadianadragan, @AllanDafoe & others https://deepmind.google/..…
-
@rohinmshah
Rohin Shah
on x
I've really liked @METR_Evals approach to rigorously evaluate the plausibility of concrete AI threat models, and say in advance how to mitigate them. I'm excited that we at @GoogleDeepMind have now made our contribution! https://twitter.com/...
-
@davlindner
David Lindner
on x
I'm very excited about @GoogleDeepMind announcing our approach to responsible scaling of frontier models. With AI moving as fast as it is, it is critical to keep assessing risks and anticipate and plan for future risks!
-
@stevesi
Steven Sinofsky
on x
This is DeepMind's new AI “Safety Framework”. One of the weird things about this, aside from the general topic of AI safety, is that the high-order bit is to focus on protecting the model weights. That seems a lot less about safety and a lot more about trade secret in practice. […
-
@_lewisho
Lewis Ho
on x
GDM's 1st step towards the ambitious ideals of responsible scaling, these being: identifying AI capabilities that pose severe risk, using evals to detect such capabilities, preparing and articulating mitigations plans, and involving externals in the process as appropriate.
-
@neelnanda5
Neel Nanda
on x
I'm really excited to see Google DeepMind's Frontier Safety Framework (similar to a Responsible Scaling Policy/Preparedness Framework) come out! Thanks to the team for all the hard work that went into writing this
-
@mikeknoop
Mike Knoop
on x
@stevesi This unfortunate trend started with the GPT-4 paper. [image]
-
@mrgunn
@mrgunn
on x
It would be interesting to compare this with @AnthropicAI's RSPs, and maybe also @OpenAI's preparedness framework, though I'm not sure how much of that team is still around.
-
@ancadianadragan
Anca Dragan
on x
Proud to share one of the first projects I've worked on since joining @GoogleDeepMind earlier this year: our Frontier Safety Framework. Let's proactively assess the potential for future risks to arise from frontier models, and get ahead of them! https://deepmind.google/...
-
@reedalbergotti
Reed Albergotti
on x
Scoop: Google DeepMind just dropped a framework for how to evaluate future AI models as they are being created, avoiding potentially dangerous capabilities: https://www.semafor.com/...
-
@manderljung
Markus Anderljung
on x
We've now got public statements from the three major frontier AI developers about what system deployments they'd consider unacceptably risky. That's a good start! Next up I'm keen to see more folks properly engage with their content, spot their flaws, and help improve them.
-
@andrewcurran_
Andrew Curran
on x
Google has designed a new Frontier Safety Framework. They will evaluate models under the new framework ‘every 6x in effective compute and for every 3 months of fine-tuning progress’. CCL stands for Critical Capability Levels. [image]
-
@allandafoe
Allan Dafoe
on x
As we push the boundaries of AI, it's critical that we stay ahead of potential risks. I'm thrilled to announce @GoogleDeepMind's Frontier Safety Framework - our approach to analyzing and mitigating future risks posed by advanced AI models. 1/N https://deepmind.google/...
-
@mihonarium
Mikhail Samin
on x
At first glance, worse than Anthropic's RSPs, the interesting things are in the Future work section
-
@ancadianadragan
Anca Dragan
on x
Leading to the Frontier Safety Framework was our dangerous capabilities evals work, expansively probing at capabilities to self-proliferate, self-reason, perform harmful cyber, and persuade. Hope it sets a new bar for pre-deployment evals! https://arxiv.org/...