DeepMind runs a double-blind evaluation of a frontier model

Google DeepMind / blogPress kit
Google DeepMind described a completed pilot on 27 August 2026 that it calls the first double-blind evaluation of a proprietary, frontier-class AI model. The model tested was Gemini Flash Lite, assessed against confidential benchmarks held by outside organisations.
The arrangement addresses benchmark contamination. When a model has seen evaluation questions during training, its scores rise without its capability changing, and the usual defence is to keep tests secret, which means the people holding them cannot run them on a model whose weights the developer will not release. Each side has something it will not hand over. The pilot ran the evaluation inside a Google Cloud Confidential Space, a hardware-backed enclave with GPU support, so evaluators never saw the weights and DeepMind never saw the prompts, with both properties verified cryptographically rather than promised.
The named partners are the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons.
The result is procedural rather than a capability finding: it establishes that an independent party can measure a closed model without either side surrendering its asset. DeepMind reports this as a pilot on one comparatively small model, so nothing yet shows the method scales to a flagship system or that external evaluators can compel access to one.