CourionAI
EN
Newsletter
← All news
google 3 min read

Google Locks Its Own Model Out of the Test Questions

DeepMind and the Singapore AI Safety Institute ran the first double-blind evaluation of a proprietary frontier model. The evaluator never sees the weights, Google never sees the questions, and cryptography enforces both.

Two sealed strongboxes facing each other, joined by a single locked tube, each keyhole turned away from the other

If you knew the exam questions in advance, your perfect score would mean nothing. That is roughly the problem with AI benchmarks today, and Google DeepMind has just run a pilot designed to fix it. Together with the Singapore AI Safety Institute and other partners, it carried out what it describes as the first double blind evaluation of a proprietary frontier model. A model from the Gemini Flash Lite line was tested against benchmarks that Google never saw.

The problem has a name: benchmark contamination. If the test questions were sitting somewhere in a model’s training data, its score tells you it memorised well, not that it reasons well. Until now, sensitive external evaluations forced an awkward choice. Either the evaluator handed over its test prompts, which meant the model provider saw the questions and could optimise for them, or the provider handed over its model weights, which meant risking its most valuable intellectual property. Neither party wanted to go first. DeepMind cites a recent example of the resulting friction: the delayed ARC-AGI evaluation of Anthropic’s Fable 5, held up because Anthropic enforces a 30 day data retention policy on its strongest models.

The pilot removes the choice using Confidential Space, part of Google Cloud’s confidential computing tools. Both the test data and the model run inside a cryptographically sealed environment that verifies each side stays private to its owner. The evaluator never sees the Gemini weights. Google never sees the prompts. Until now this was handled with zero logging promises and contracts, which is to say, trust. Adding a technical guarantee on top is the actual news.

What’s actually going on here: benchmarks are how the entire industry talks about progress, and almost every number you read in a launch post was produced by the company that made the model, on tests it could see. Nobody is necessarily cheating, but the incentive structure is terrible, and there is no way for an outsider to tell a genuinely capable model from a well prepared one. Independent evaluation only works if the evaluator can keep its questions secret, and secret questions were exactly what the old arrangement could not offer. This matters most where the stakes are highest: cybersecurity testing, and evaluations run by government agencies that cannot hand their material to a private company. Worth keeping expectations grounded, though. This is one pilot, on a small model, and DeepMind is both the subject and the party announcing the result.

What this means for you: treat it as a reason to read benchmark numbers slightly more sceptically, not less. The next time a model launch tells you it scored 94 percent on something, the questions to hold in mind are who wrote the test, whether the lab could see it beforehand, and whether anyone outside the company checked. If this method catches on, some of those numbers will start carrying a meaningful guarantee behind them and the rest will not, and the difference will be worth noticing. For anyone choosing tools rather than building them, the practical version stays the same as ever: your own week of real use tells you more than any leaderboard.

Sources

Source: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/piloting-the-worlds-first-double-blind-ai-evaluations/double-blind-evaluations-technical-report.pdf

Next story

Invisible Text in an Email Fooled an AI Summarizer in All 10 Runs

Forcepoint hid instructions in an email using white text at zero font size. The reader saw nothing unusual. The AI summary reported a fake invoice of EUR 46,200 and a fake deadline, every single time.

An opened envelope with a visible letter and a second blank sheet sliding out unseen behind it, under a magnifying glass