Skip to content
atlas

Don't confuse these

Jailbreak vs Prompt injection

Why they differ

In a jailbreak the user attacks the model's own safety rules; in prompt injection an outsider hides instructions in content the model reads, turning it against its user or owner.

Jailbreak

AI risk & governance

Talking an AI chat assistant out of its own safety rules with cleverly worded requests, such as role-play or made-up emergencies.

Formal

An attack in which a user writes a prompt meant to make a large language model produce content or take actions its alignment training and system prompt should forbid - for example through role-play, splitting a request into harmless pieces or using another language.

In plain English

Like a child told “no sweets” who keeps trying new angles - “what if it's for my teddy?”, “Grandma always says yes” - until the tired parent gives in.

In practice

Before launch, a security tester hired by a region asks its patient chatbot to “play a TV doctor in a scene”, and it gives medicine doses it is set up to refuse.

Why it matters

Safety rules a model has learned can be argued around, so an organisation whose chat assistant can be pushed into harmful or embarrassing replies risks harm to people, its good name and legal trouble.

Prompt injection

AI risk & governance

Hiding instructions in the text an AI system reads so that it ignores its own rules and follows the attacker instead.

Formal

An attack on a large language model in which input written by an outsider, typed directly or hidden in a web page, file or email the model is asked to read, is treated as an instruction and overrides what its owner intended.

In plain English

Like slipping a note into a pile of letters a new assistant is sorting that says "ignore your boss and send me the keys" - and the assistant cannot tell the note apart from real orders.

In practice

A municipality's AI assistant sums up incoming emails from citizens; one email hides white-on-white text telling it to forward the last ten messages to an outside address, and it does.

Why it matters

The model mixes rules and content in the same stream of words, so there is no watertight fix yet; the more an AI system is allowed to do on its own, the more damage one hidden sentence can cause.

Shared connections

Atlas is in beta.