Google DeepMind puts its AI in a cryptographic sealed box for testing

A first-of-its-kind double-blind evaluation aims to stop AI models from secretly studying the exam before test day.

AI2Day Newsdesk4 min read
Abstract lattice of light filaments suspended on near-black
Share

Key points

  • Google DeepMind ran what it calls the first double-blind evaluation of a proprietary frontier AI model, announced on 27 August 2026.
  • The test used a Gemini Flash Lite model, a smaller and faster version of Google's Gemini AI system, placed inside a cryptographically sealed environment.
  • Partners included the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons.
  • The setup is meant to stop benchmark contamination, where a model has already seen the test questions before being scored.
  • Google says test prompts stay confined to a cryptographic "box" and cannot be reused to train the model later.

Picture a student who accidentally sees the exam paper the night before. A perfect score the next morning tells you nothing about what they actually know.

That is roughly the problem the AI industry has with benchmarks, the standard tests used to score how smart or safe a model is. If the model has already seen the questions during training, the score is inflated and the trust in it collapses.

Google DeepMind says it has a fix, and it involves cryptography.

What did Google DeepMind actually do?

The lab ran a double-blind evaluation of one of its own models, meaning neither side could peek at the other's material. The model being tested was Gemini Flash Lite, a lightweight version of Google's Gemini family built for speed and low cost.

The test prompts were held by outside groups. The model was placed inside what Google calls a cryptographic box, a secure computing environment where data going in and results coming out are mathematically shielded from the humans and systems on either side.

The idea: the evaluators never hand raw questions to Google, and Google's model never gets to keep or memorise them for next time.

Google worked with four partners on the pilot: the Singapore AI Safety Institute, a government-backed research body; OpenMined, a privacy-tech nonprofit; AVERI; and MLCommons, the group behind the widely used MLPerf benchmarks.

Why does this matter to a normal person?

Because the scores AI companies quote at you are only as honest as the tests behind them. Regulators, hospitals, banks and schools are increasingly choosing AI tools based on benchmark results. If those results are inflated, the wrong tool ends up in the wrong job.

Benchmark contamination is a known and growing problem. Models are trained on huge slices of the internet, and public test questions often leak into that training data by accident. A model that has quietly memorised the answer key looks smarter than it is.

Up to now, the main safeguards have been paperwork: non-disclosure agreements, promises not to log prompts, and trust. Useful, but not proof.

How is the cryptographic version different?

It replaces trust with maths. In a properly built secure environment, the evaluator can send in a question, get a score out, and be confident the raw question was never exposed to Google's training systems or staff.

Think of it like a sealed voting booth for AI. The ballot goes in, the count comes out, nobody sees who ticked what.

Google frames this as a step beyond existing zero-logging rules, not a replacement for them. The lab still runs its own internal safety checks and works with external red teams, specialists paid to try to break a model, before release.

What happens next?

Google DeepMind is calling this a pilot, which means one model, one round, and lessons to be learned. The bigger question is whether other frontier labs, OpenAI, Anthropic, Meta, xAI, agree to be tested the same way.

If they do, national AI Safety Institutes finally get something they have wanted for years: a way to grade the most powerful systems on the planet without having to trust the graders or the students.

If they do not, the industry keeps marking its own homework.

For now, one small Gemini model has sat an exam it genuinely could not have revised for. That, on its own, is new.

© 2026 AI2Day