Red teaming
Deliberately attacking your own system to find out how it breaks before someone else does.
Red teaming comes from military and security practice: one group plays the attacker and tries to get through the defences, so the defenders learn where the holes are. In AI, it means people whose job is to make a model misbehave. They try to talk it into giving dangerous instructions, trick it with hidden text, or push it into situations its designers never considered.
Labs run red teaming before releasing a model and publish results as evidence that a system is safe enough to ship. Worth reading those claims carefully. Red teaming shows what testers managed to find, not what exists, so it can prove a weakness is there but never prove one is absent. It also tends to be strongest against attacks people already know about, which is why new attack styles keep succeeding against models that passed their tests.
-
Watching Its Own Model Now Costs OpenAI 20 Percent Extra, and It Paused a Training Run to Do It
-
Claude Code Turns On Auto Mode by Default, and the Safety Numbers Are Not What You Would Guess
-
Anthropic says Claude found real weaknesses in two encryption algorithms
-
OpenAI Built an AI Whose Only Job Is to Attack Its Other AIs