Red-teaming
Structured adversarial testing of an AI system to find flaws, harmful behaviours and vulnerabilities before attackers do.
AI red-teaming is a structured testing exercise in which a team probes an AI system, often adversarially, to find flaws, vulnerabilities and harmful or unexpected behaviours. NIST's Generative AI Profile describes it as an evolving practice for identifying potential adverse behaviour or outcomes of a model or system, how they could occur, and for stress-testing safeguards. For agents, red-teaming should target actions as well as words, for example trying to make the agent pay a new payee, exceed a limit or skip an approval.
Agent Minute explains this term on 21 February 2027.
Related terms
Model validationEvaluating whether a model performs as expected, including its reliability and its limitations.Prompt injectionAn attack that alters what a model receives so that it follows the attacker's instructions instead of its user's.JailbreakA prompting attack designed to get a model to bypass its safety rules or restrictions.Excessive agencyGiving an LLM application more functions, permissions or autonomy than needed, so bad outputs cause harmful actions.