/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Google launches a pilot of double-blind AI evaluations, keeping external evaluations in a cryptographic “box” to stop benchmark contamination and protect IP

Building trust in proprietary model benchmarks using cryptographically secure environments  —  Imagine a student is set to take a high-stakes exam.

Google DeepMind

Discussion

  • @googledeepmind @googledeepmind on x
    In an industry first, we're piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and
  • @jayanayak21 Jaya Nayak on x
    🚨 What if AI safety tests could be compromised before they even begin? 🧐👀 Google is piloting an industry-first “double-blind evaluation system for frontier AI”. • Test prompts remain hidden • Model weights remain private • External evaluators can test safety and
  • @zswaff Zack Swafford on x
    This is an annoying unsolved problem currently for the community: sharing benchmarks = they get into training data = benchmark becomes near useless. Keeping the benchmark test data and model weights private while still providing verifiable results is valuable.
  • @bakkermichiel Michiel Bakker on x
    Super important work from @AVERIorg @openminedorg , @GoogleDeepMind and @MLCommons that deserves much more attention, I think! It's the first double-blind eval of a closed-weight model. A quick post on why I think this is important. Normally, there are only two options to
  • @iamtrask @iamtrask on x
    If all AI companies and AI eval organizations used double-blind evals, test set contamination in training data would disappear overnight, and confidence in AI eval veracity would rise across all subject areas.
  • @0xj4yd3v Jay Dev on x
    Double-blind evals only matter if every lab submits to the same third party. Replies are asking who verifies the enclave - the harder problem is getting the other labs to agree to this at all.
  • @koenvanderveen Koen van der Veen on x
    Today, we released our work on double blind eval of frontier models using PySyft. To the best of my knowledge this is the first of its kind. Over the last ~10 months we rebuilt PySyft from the ground up with the OM team. Still a lot to do but very happy with these results!
  • @averiorg @averiorg on x
    Today we're announcing a historic milestone: The first ever double-blind evaluation of a proprietary language model. This was made possible by a unique collaboration between AVERI, @GoogleDeepMind, @OpenMinedOrg, and @MLCommons. We tested Gemini 2.5 Flash-Lite using
  • @webthreeai @webthreeai on x
    First time a frontier lab is doing double-blind evals. Google says in this pilot, neither the test prompts nor the model weights are revealed to the external evaluators aiming for private, robust, trustworthy safety checks. What should be measured first?
  • @jastephx Jason Stephen on x
    Eval sets stay confidential, while model weights stay private, win-win!
  • @solomonmg Sol Messing on x
    Thrilled to announce a step toward fixing benchmark contamination—the first successful test of our confidential evaluation framework “double blind evals” on a frontier class proprietary model with a non-public evaluation set.
  • @openminedorg @openminedorg on x
    After nearly a decade of research and development including contributions from over 400 contributors from around the world, we are pleased to announce that PySyft has been used by @GoogleDeepMind, @AVERIorg, Singapore AISI, and @MLCommons to facilitate the world's first
  • Helen King Helen King on linkedin
    When evaluating AI models, external testing is vital in providing further perspectives and an independent view on our capabilities and progress. …