← Glossary Model
GPT-Red
An internal OpenAI model built to attack its own systems and find security holes.
GPT-Red is a model OpenAI trained specifically to probe its other systems for weaknesses, acting like an automated attacker so problems can be caught and fixed before real attackers find them. Reports say it finds flaws more often than human experts.
It is a striking example of AI being turned on itself for defence, and of why security is becoming central as models start acting on people’s behalf.
Mentioned in