Google DeepMind announced on August 27 it is piloting what it calls the world's first double-blind evaluation of a frontier AI model, sealing confidential benchmark questions inside a cryptographic box so the model cannot study them before the test S¹. The pilot runs a Gemini Flash Lite model against hidden benchmarks with four external partners, including the Singapore AI Safety Institute and MLCommons. What the announcement does not reveal is whether that box actually works, or whether anyone outside Google has verified it.
My read: This is the first time a major lab has publicly admitted, by trying to fix it, that the standard evaluation process has a structural hole. The problem is real: if a model trains on its own benchmark questions, the scores are theatre. But a cryptographic box designed by the same company that built the model is a bit like a student bringing their own proctor to the exam. I want to see the protocol audited by someone who did not sign DeepMind's NDA before I trust the "double-blind" label. The partners are credible names, but the blog post names them without describing what they actually do in the pilot.
Why benchmark contamination is the quiet crisis in AI
Every AI model gets scored on benchmarks: standardised test sets that measure reasoning, coding, safety, and other capabilities. If the model has already seen those questions during training, it scores high without being genuinely capable. Researchers call this contamination, and it is one of the quietest problems in the field.
DeepMind says it has relied on zero-logging rules and contractual safeguards to keep external test prompts confidential until now S¹. In plain terms: they have been depending on promises and legal agreements. A contract is only as strong as the incentive to break it, and the incentive to top a benchmark leaderboard is enormous. High scores drive attention and revenue.
The new pilot adds a technical layer on top of the legal one. The evaluation keeps external test questions confined to a cryptographic environment where models cannot access them later to optimise performance S¹. Think of a sealed envelope: the tester writes the questions, locks them in a box the model cannot open, and the model answers without ever seeing the questions during training.
DeepMind's crypto box applies the same principle to software: verification that the test questions have not leaked into the training data.
Who is in the room
DeepMind is partnering with four organisations: the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons S¹. MLCommons runs the MLPerf benchmark suite that the chip industry uses to compare hardware. OpenMined works on privacy-preserving machine learning. The Singapore AISI has been building agent-testing infrastructure: a GitHub repo shows bilateral agent testing by the Korea and Singapore AI Safety Institutes, created in January 2026 P⁵.
The model under test is Gemini Flash Lite, the efficiency tier of Google's Gemini family. Google introduced the Gemini 3.5 and 3.6 Flash lineup on July 21, pitching the Flash Lite tier at developers building AI agents at scale P⁶. Testing the smallest model first makes sense: if the cryptographic setup breaks something, you want to find out on the cheapest model, not the flagship.
What to do about it
If you build or buy AI systems, the contamination problem is not abstract. Consider a hospital's AI procurement team comparing models for clinical decision support. They read benchmark scores showing one model at 92% accuracy on medical reasoning and another at 87%. If the first model trained on the benchmark itself, that 92% is a mirage. The hospital buys the wrong model, and the gap shows up later in patient outcomes, not in a scorecard.
The practical takeaway: when a vendor cites benchmark numbers, ask whether those benchmarks were run under contamination controls. If the answer is "we have a contract," that is weaker than "we have a cryptographic guarantee." DeepMind's pilot is the first time a major lab has publicly attempted the latter. Until results are published, treat any benchmark score from any lab as a best-case claim, not a verified measurement.
One thing you can do this week: check whether your organisation's model evaluation process includes any technical safeguard against training-data contamination, or just a policy document. The gap between those two is where the risk lives.
What we don't know yet
The blog post does not include evaluation results, a completion date, or any technical detail about how the cryptographic box works S¹. The "world's first" claim is Google DeepMind's own, and no independent party has verified it. The partners are named, but their specific roles, whether they designed the protocol or audited the results, are not described.
The cryptographic method has not been independently audited, based on the evidence available. A box is only as trustworthy as the protocol it runs on, and the protocol is not public.
No timeline or completion date has been disclosed. The next signal: if MLCommons or the Singapore AISI publishes its own account of the pilot, that would confirm whether the "double-blind" label survives outside scrutiny. We will check the claim against it when it appears.
When the first partner-published results land, we will break down what the box actually proved. Subscribe to catch that.
Sources: S1 — Piloting the world's first double-blind AI evaluations · S2 — Piloting the world's first double-blind AI evaluations - Google DeepMi · P3 — Towards Conversational Medical AI with Eyes, Ears and a Voice · P4 — google-deepmind/tips · P5 — sgaisi/kr-sg-aisi-agent-testing · P6 — 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
More from Not A Tech Guy
- AI coding agent defense cuts malware severity 83%
- WebMCP-Phalanx blocks 80 of 80 prompt injection attacks in browser agents
- StepGuard blocks AI agent attacks 77% before they run
Generated from an audited evidence pack with primary-source research. Social-media items are discussion signals, not verified facts. Nothing here is financial, legal or medical advice.