AI risks and safety

Jailbreak

A prompting attack designed to get a model to bypass its safety rules or restrictions.

A jailbreak is a prompting attack designed to make a generative AI model bypass the safety restrictions or usage rules placed on it by its developer or operator, for example through role play, obfuscation or long chains of instructions. NIST's taxonomy of adversarial machine learning treats jailbreaks as a category of attacks on generative models. Because jailbreaks keep evolving, controls that matter for money must not depend on the model refusing.

Agent Minute explains this term on 15 November 2026.

Related terms