Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Google DeepMind announces the world's first double-blind evaluation of a proprietary frontier AI model, using cryptographic environments to prevent benchmark contamination. The pilot tests a Gemini Flash Lite model with external partners against confidential benchmarks in a privacy-preserving environment, enhancing evaluation integrity and trust in model assessments.
From the source
Today, we’re introducing the world’s first double-blind evaluation of a proprietary, frontier class AI model, which keeps external evaluations confined to a cryptographic “box” where they can’t be used by models later to optimize performance ahead of testing. We're partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons , to test a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving environment , increasing evaluation integrity.
deepmind.google