TL;DR
Google DeepMind has piloted what it describes as the first double-blind assessment of a commercial frontier-class system. A Gemini Flash Lite system was tested against confidential benchmarks inside a cryptographically sealed environment, with MLCommons, AVERI, OpenMined and Singapore’s AI Safety Institute taking part.
The contamination problem
Benchmark scores carry weight with regulators, buyers and researchers, and they are only meaningful if the model has not previously seen the questions. Once test material leaks into training data, a high score measures recall rather than capability, and there is no reliable way to tell the difference after the fact.
Until now, serious external testing forced an awkward trade. Either the evaluator handed its prompts to the lab, which then holds the exam paper, or the lab handed over model weights, exposing the asset the whole business rests on. Confidentiality was maintained by contract and by zero-logging commitments — arrangements that depend on trust rather than proof.
What changed technically
The pilot runs the evaluation inside Confidential Space, part of Google Cloud’s confidential computing range, which produces cryptographic attestation that neither side saw the other’s material. The testers get no access to the weights; Google gets no sight of the prompts. Neither party has to be believed, because the environment can be verified.
That distinction matters most where the stakes are highest. Cybersecurity evaluations and government-run testing are precisely the cases where a lab’s assurance that it did not look is worth least, and where the evaluating body may be legally unable to share its material at all.
Looking forward
This lands the same week the UK AI Security Institute published its own work on evaluation efficiency, and the pairing is the story. National safety institutes, DeepMind’s London base among the labs they scrutinise, have spent two years running pre-deployment testing under negotiated access arrangements whose terms are largely invisible from outside.
Cryptographic attestation offers a way to make those arrangements checkable rather than confidential, which would let a testing body publish a result that stands on its own. It also lowers the barrier for institutes in jurisdictions that cannot lawfully hand sensitive evaluation material to an American company.
The pilot involves one lightweight model and four partners, so nothing about frontier oversight has yet changed. Whether it becomes standard practice depends on whether other labs submit their largest models to the same conditions — the ones where a contaminated benchmark would matter most.