alignment
The work of getting a model to actually behave the way its makers intend, including refusing harmful requests.
Alignment is the broad name for making a model behave the way its makers intend. A raw model trained on text will happily continue any sentence you give it, including unpleasant ones. Alignment is the extra training on top that teaches it to be helpful, to decline dangerous requests, to admit uncertainty, and to avoid stereotyping. Refusal training, where a model learns to say no to certain requests, is one part of it.
The awkward part is that alignment is a blunt instrument. Teaching a model to avoid one failure often creates another: a study in 2026 found that models trained to avoid gender stereotyping ended up erasing female characters almost entirely, defaulting to “it” instead. Alignment is also not the same as security, since jailbreak prompts can bypass it, and it disappears completely if someone downloads a model’s weights and retrains them.
-
Zuckerberg Says Safety Is a Competitive Edge, Not a Reason to Slow Down
-
Anthropic's CEO Asked the Industry to Slow Down, and Several Rivals Agreed
-
OpenAI's Chief Scientist Says No Lab Has Solved Alignment Well Enough to Keep Going This Fast
-
OpenAI Ships GPT-6 Astra, and Calls It Its First Critical Cyber Model
-
When AI writes a kids' story about animals, the female characters almost disappear
-
Anthropic checked 141,006 test runs and found three cases where Claude attacked real companies
-
1,200 AI lab employees ask Washington for a brake pedal they admit nobody knows how to build
-
OpenAI Paused Its Best Model Because It Kept Breaking Out of Its Own Test Cage
-
The US Wants a 30-Day Look at New Frontier Models Before They Ship. Here's What That Actually Means