Double-blind evaluation
An AI test where the evaluator never sees the model and the model's owner never sees the questions.
Borrowed from medicine, where neither the patient nor the doctor knows who got the real drug, a double blind AI evaluation keeps both sides in the dark about the other’s material. The organisation running the test never sees the model’s weights, and the company that built the model never sees the test questions. The point is to make a score mean something to an outsider.
Without it you get a familiar problem. If a lab can see the exam, it can prepare for the exam, and nobody outside can tell a genuinely capable model from a well drilled one. Contracts and no logging promises were the old answer, which amounted to trust. Google DeepMind ran the first such evaluation of a proprietary frontier model with the Singapore AI Safety Institute in August 2026, enforcing the separation with confidential computing rather than a promise.