AI risk & governance
Talking an AI chat assistant out of its own safety rules with cleverly worded requests, such as role-play or made-up emergencies.
Formal
An attack in which a user writes a prompt meant to make a large language model produce content or take actions its alignment training and system prompt should forbid - for example through role-play, splitting a request into harmless pieces or using another language.
In plain English
Like a child told “no sweets” who keeps trying new angles - “what if it's for my teddy?”, “Grandma always says yes” - until the tired parent gives in.
In practice
Before launch, a security tester hired by a region asks its patient chatbot to “play a TV doctor in a scene”, and it gives medicine doses it is set up to refuse.
Why it matters
Safety rules a model has learned can be argued around, so an organisation whose chat assistant can be pushed into harmful or embarrassing replies risks harm to people, its good name and legal trouble.