← Glossary Term
CyberGym
A benchmark that measures how well AI models find and exploit real software vulnerabilities.
CyberGym is a benchmark for offensive security skill. It puts a model in front of real codebases with known vulnerabilities and scores whether it can find the weakness and produce something that actually triggers it, rather than just describing the bug in the abstract.
It has become one of the headline numbers labs cite when launching security-focused models, which is a slightly uncomfortable thing to be competing over. The same capability that lets a model find a flaw before attackers do is the capability that helps attackers find it first, which is why several labs restrict access to their strongest models in this area.
Mentioned in
-
GLM-5.3 Got Much Better at Breaking Software, So Z.ai Is Holding the Weights Back
-
An open model is now months, not years, behind the frontier. Its safety testing is not.
-
Microsoft says it will stop chasing the frontier and build small specialist models instead
-
1,200 AI lab employees ask Washington for a brake pedal they admit nobody knows how to build
-
Hugging Face published the full forensic timeline of the AI agent that broke into its systems
-
Microsoft's new security model is small on purpose, and that is the interesting part