Jailbreak
A prompting attack designed to get a model to bypass its safety rules or restrictions.
A jailbreak is a prompting attack designed to make a generative AI model bypass the safety restrictions or usage rules placed on it by its developer or operator, for example through role play, obfuscation or long chains of instructions. NIST's taxonomy of adversarial machine learning treats jailbreaks as a category of attacks on generative models. Because jailbreaks keep evolving, controls that matter for money must not depend on the model refusing.
Agent Minute explains this term on 15 November 2026.
Related terms
Prompt injectionAn attack that alters what a model receives so that it follows the attacker's instructions instead of its user's.System promptThe standing instructions an operator gives a model before any user input, describing its role, rules and tools.Red-teamingStructured adversarial testing of an AI system to find flaws, harmful behaviours and vulnerabilities before attackers do.Excessive agencyGiving an LLM application more functions, permissions or autonomy than needed, so bad outputs cause harmful actions.