Google launches a pilot of double-blind AI evaluations, keeping external evaluations in a cryptographic “box” to stop benchmark contamination and protect IP
Building trust in proprietary model benchmarks using cryptographically secure environments — Imagine a student is set to take a high-stakes exam.
Google DeepMind
Related Coverage
- Google found a way to test Gemini without seeing the questions The New Stack · Amanda Caswell
- DeepMind Launches First Double-Blind AI Model Evaluation Blockchain.News · Luisa Crawford
- Google DeepMind pilots double-blind AI evaluations without sharing prompts or model weights RuntimeWire
- AI benchmarks have a trust problem and Google wants to fix it The Decoder · Maximilian Schreiner
Discussion
-
@googledeepmind
@googledeepmind
on x
In an industry first, we're piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and
-
@jayanayak21
Jaya Nayak
on x
🚨 What if AI safety tests could be compromised before they even begin? 🧐👀 Google is piloting an industry-first “double-blind evaluation system for frontier AI”. • Test prompts remain hidden • Model weights remain private • External evaluators can test safety and
-
@zswaff
Zack Swafford
on x
This is an annoying unsolved problem currently for the community: sharing benchmarks = they get into training data = benchmark becomes near useless. Keeping the benchmark test data and model weights private while still providing verifiable results is valuable.
-
@bakkermichiel
Michiel Bakker
on x
Super important work from @AVERIorg @openminedorg , @GoogleDeepMind and @MLCommons that deserves much more attention, I think! It's the first double-blind eval of a closed-weight model. A quick post on why I think this is important. Normally, there are only two options to
-
@iamtrask
@iamtrask
on x
If all AI companies and AI eval organizations used double-blind evals, test set contamination in training data would disappear overnight, and confidence in AI eval veracity would rise across all subject areas.
-
@0xj4yd3v
Jay Dev
on x
Double-blind evals only matter if every lab submits to the same third party. Replies are asking who verifies the enclave - the harder problem is getting the other labs to agree to this at all.
-
@koenvanderveen
Koen van der Veen
on x
Today, we released our work on double blind eval of frontier models using PySyft. To the best of my knowledge this is the first of its kind. Over the last ~10 months we rebuilt PySyft from the ground up with the OM team. Still a lot to do but very happy with these results!
-
@averiorg
@averiorg
on x
Today we're announcing a historic milestone: The first ever double-blind evaluation of a proprietary language model. This was made possible by a unique collaboration between AVERI, @GoogleDeepMind, @OpenMinedOrg, and @MLCommons. We tested Gemini 2.5 Flash-Lite using
-
@webthreeai
@webthreeai
on x
First time a frontier lab is doing double-blind evals. Google says in this pilot, neither the test prompts nor the model weights are revealed to the external evaluators aiming for private, robust, trustworthy safety checks. What should be measured first?
-
@jastephx
Jason Stephen
on x
Eval sets stay confidential, while model weights stay private, win-win!
-
@solomonmg
Sol Messing
on x
Thrilled to announce a step toward fixing benchmark contamination—the first successful test of our confidential evaluation framework “double blind evals” on a frontier class proprietary model with a non-public evaluation set.
-
@openminedorg
@openminedorg
on x
After nearly a decade of research and development including contributions from over 400 contributors from around the world, we are pleased to announce that PySyft has been used by @GoogleDeepMind, @AVERIorg, Singapore AISI, and @MLCommons to facilitate the world's first
-
Helen King
Helen King
on linkedin
When evaluating AI models, external testing is vital in providing further perspectives and an independent view on our capabilities and progress. …