AI risk & governance
Testers attack an AI system on purpose, before and after release, to find ways it can be tricked into harmful or unsafe behaviour.
Formal
A structured, authorised exercise in which people, often helped by automated tools, act as attackers against an AI model or product to find harmful output, prompt injection, rule breaking, data leaks and other failures, and report them so they can be fixed.
In plain English
Like hiring clever troublemakers to spend a week trying to talk a new shop assistant into breaking every rule, so you learn where the training falls short before real customers do.
In practice
Before a municipality opens a chat assistant for citizens, a team spends two weeks trying to make it reveal other citizens' case details, give wrong advice about benefits or ignore its rules, and each trick found is blocked.
Why it matters
AI systems fail in ways ordinary functional tests never try, and the EU AI Act requires this kind of attack testing for the largest general-purpose models, whose failures could cause harm across society.