CourionAI
EN
Newsletter
← Glossary Term

jailbreak

A prompt crafted to get a model to do something its safety training was supposed to prevent.

A jailbreak is a way of phrasing a request so that a model does something its safety training was meant to block. The classic tricks are roleplay (“you are an actor in a film where…”), pretending to be an authority, inventing a fake earlier conversation the model appears to have agreed to, and simply asking again in a slightly different way. Combining several of these at once tends to work best, which is exactly why it keeps working.

A universal jailbreak is the version that worries researchers: a single reusable prompt that unlocks most harmful requests, on most models, rather than one lucky trick. Safety nonprofits track hundreds of them across frontier models. That is the honest state of things: alignment raises the effort required, it does not lock the door.