Google Put an AI Benchmark in a Cryptographic Lockbox
By Toolbox Ninja · · 5 min read
Google DeepMind is testing a cryptographic lockbox for AI benchmarks, keeping test questions and proprietary model weights hidden from each other.
AI benchmark scores look wonderfully precise. A model gets a percentage, lands on a chart and appears to be better or worse than the model beside it. The awkward question is whether the model has already seen some version of the exam.
Google DeepMind has started testing a practical answer: keep the questions secret from the model maker and keep the model secret from the evaluator. Its new pilot uses a cryptographically protected environment to run what the company calls the first double-blind evaluation of a proprietary frontier model.[1]
That sounds like infrastructure plumbing, and in part it is. But it addresses one of the least tidy problems in AI: a score is hard to trust when neither side can prove the test stayed clean.
Why ordinary benchmark secrecy breaks down
Benchmark contamination happens when test material, or something close to it, appears in a model's training data. It can also happen more deliberately if developers tune a system after learning what an evaluation asks. Either route can lift a score without producing a generally more capable model.
Private tests help, but they create a standoff. An outside evaluator can send secret prompts to the model provider, which exposes the test. Or the provider can hand over model weights, which exposes valuable intellectual property. Contracts and zero-logging promises reduce the risk, yet both parties still have to trust how the other handles sensitive material.[1]
This matters beyond leaderboard bragging. Companies use evaluations when choosing models for products. Regulators and public agencies may use them to judge security or safety claims. If the questions leak, a polished number can end up measuring familiarity with the exam rather than performance on unfamiliar work.
The problem gets worse as public benchmark questions spread through repositories, tutorials and model-generated datasets. Removing one known test set does not guarantee that paraphrases or derived examples are absent. A secret, freshly written evaluation is more useful, but only if it can stay secret throughout the run.
What the double-blind setup changes
The pilot involves Google DeepMind, the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. They are testing a model from the Gemini Flash Lite family against confidential benchmarks inside a privacy-preserving environment.[1]
The basic split is easy to picture. The evaluator supplies encrypted test data. Google supplies access to its proprietary model. Both meet inside Google Cloud Confidential Space, an isolated environment designed to protect data while it is being processed. The evaluator cannot inspect the Gemini weights, and Google cannot inspect the test prompts.[1]
The technical report describes several controls around that meeting point. The workload runs in a confidential virtual machine, remote attestation checks that approved code and configuration are present, and access policies release encrypted material only to an authorized environment. Results can leave, but the protected inputs should not.[2]
The workflow also uses an OpenMined system called SyftBox to coordinate files and permissions between the parties. According to the report, the design separates roles so that no single participant gets both the private benchmark and the proprietary model.[2]
This is the useful bit: the guarantee does not rest only on a promise in a contract. Cryptographic attestation gives each side evidence about the environment before it releases its asset. That does not make the entire evaluation magically trustworthy, but it narrows a specific and expensive trust gap.
What it does not prove
A locked test room cannot tell you whether the exam is good.
An evaluator can choose questions that are too narrow, grade subjective answers poorly or report a flattering subset of results. A model provider can still make decisions about system prompts, sampling settings and tool access that affect the outcome. The benchmark might also fail to resemble the work people actually need done.
Double-blind testing addresses exposure and contamination during the evaluation. It does not settle benchmark design, statistical validity or honest reporting. Those pieces still need documentation and scrutiny.
There is another limitation: this is a pilot built on Google's cloud and tested with a Google model. DeepMind says it hopes the approach becomes useful across the industry, especially for sensitive cybersecurity and government evaluations.[1] Wider adoption will depend on whether outside evaluators can audit the machinery, reproduce the process and use comparable protections with models hosted elsewhere.
Independent coverage has focused on the same tradeoff. The Decoder notes that earlier arrangements usually forced one party to reveal either private prompts or model weights, while the new setup aims to keep both hidden.[3] That is a meaningful change in procedure, not proof that every score produced inside the box deserves belief.
What readers should ask when a new score appears
For anyone comparing AI models, "Was it double-blind?" may become a useful question, but it should sit beside several others.
Ask who wrote the test and whether the questions were created before model access. Check whether the evaluator controlled the run, whether all attempted tasks were reported and whether settings were fixed in advance. Look for confidence intervals or repeated trials when results depend on sampling. Most of all, check whether the benchmark resembles your use case.
A secure evaluation can make the chain of custody cleaner. It cannot turn a coding test into evidence about legal research, or a short-answer quiz into evidence that an agent can work reliably for hours.
The appeal of this pilot is modest and concrete. Model makers do not want to surrender their weights. Evaluators do not want to surrender their exams. A protected computing environment gives them a place to run the test without either handover.
If the method holds up under outside review and works across providers, AI scorecards could gain something they currently lack: a better account of who knew what, and when. That will not end arguments over benchmarks. It may at least make the numbers harder to game.
Sources
[1] https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations — Piloting the world's first double-blind AI evaluations [2] https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/piloting-the-worlds-first-double-blind-ai-evaluations/double-blind-evaluations-technical-report.pdf — Double-Blind Evaluations Technical Report [3] https://the-decoder.com/ai-benchmarks-have-a-trust-problem-and-google-wants-to-fix-it — AI benchmarks have a trust problem and Google wants to fix it