CourionAI
EN
Newsletter
← All news
ai-safety 2 min read

Nobody Passed: The First Report Card on How AI Labs Police Their Own Systems

Guidelight graded Anthropic, OpenAI, Google, xAI and Meta on six basic internal control practices. The best score was a C plus, Meta got an F, and the assessment used only public documents.

A clipboard report card with checkmarks and grade letters next to a magnifying glass over server racks

A nonprofit called Guidelight published its first assessment of how AI companies control the AI systems they run internally, and the headline finding is blunt: none of them fully apply even the basic measures. Anthropic and OpenAI lead with a C plus. Google follows with a D plus, alongside a detailed roadmap for improvement. xAI scores a D minus and Meta an F.

Guidelight checked six practices, and the list is less exotic than you might expect. Does the company log what its internal AI systems are doing? Are risky actions gated behind a review step? Is there an emergency stop, known in the field as circuit breaking? Is there a plan for containing a model that turns out to be misaligned, meaning one pursuing goals other than the ones intended? The pattern across all five companies was consistent: they are best at noticing when something has gone wrong, and worst at preventing it or containing it afterwards. Guidelight was founded by two former OpenAI safety leads, Page Hedley and Steven Adler, and drew only on public material such as system cards, safety reports and blog posts.

That last detail cuts both ways, and it is the fair caveat here. Grading a company on what it has chosen to publish rewards transparency rather than actual safety. A lab with strong internal controls and a habit of saying nothing would score badly, while a lab that writes long blog posts about its process would score well. Guidelight is upfront about this, and it is arguably the point: the group wants public documentation to become the norm precisely so that outside assessment becomes possible. The reason any of this matters now is that labs increasingly run their most capable models on themselves. Anthropic has said Claude writes most of the code in its production systems, sometimes through agents running continuously without a human watching each step. The company grading itself best still only managed a C plus.

What this means for you: for most people, nothing changes this week. The value of a report like this is as a scoreboard you can check again in six months, which is exactly what Guidelight intends. If you are choosing an AI provider for work, the six criteria are a reasonable checklist to raise with a vendor, and the honest answer to “can you show me your internal controls” tells you a lot. And if you follow the safety debate, note the shape of the finding: detection is the easy part, prevention and containment are the hard parts, and that is true for ordinary company IT security too.

Sources

Source: https://guidelight.ai/blog/control-assessment-august-2026

Next story

OpenAI Wants to Catch Misuse Without Ever Seeing Your Data

OpenAI previewed Private Safety Processing, a system meant to spot abuse patterns across several conversations while keeping zero data retention for business customers. Anthropic still requires 30 days of logs for its strongest models.

A padlock shield floating above a cloud with fading data stream lines