Google DeepMind Makes History in AI Benchmark Evaluation
Google DeepMind has made history by conducting the world’s first double-blind evaluation of a proprietary frontier-class AI model. This pilot, conducted with several partners, tested Gemini 2.5 Flash Lite against confidential benchmarks using a cryptographic hardware enclave to seal both the model weights and evaluation prompts inside a single secure environment.
The test was done in partnership with Singapore’s AI Safety Institute (AISI), OpenMined - a nonprofit focused on privacy-computing -, MLCommons, which is an industry consortium for benchmarking machine learning algorithms, and AVERI, a company that specializes in evaluation. The goal of this double-blind process is to ensure the integrity of benchmarks used to evaluate AI models.