Prompt injection
An attack that alters what a model receives so that it follows the attacker's instructions instead of its user's.
Prompt injection is an attack in which input to a generative AI system is crafted or modified so that the system behaves in ways its operator or user did not intend, such as ignoring its instructions, revealing data or calling tools. NIST's Generative AI Profile distinguishes direct prompt injection, typed by the attacker, from indirect prompt injection, hidden in content the system retrieves. For agents with tools, a successful injection can turn into an unauthorised action, so the defences belong outside the model as well as inside it.
Agent Minute explains this term on 18 October 2026.
Related terms
Indirect prompt injectionPrompt injection hidden in content an AI system retrieves, such as a web page, document or email.JailbreakA prompting attack designed to get a model to bypass its safety rules or restrictions.Excessive agencyGiving an LLM application more functions, permissions or autonomy than needed, so bad outputs cause harmful actions.System promptThe standing instructions an operator gives a model before any user input, describing its role, rules and tools.