AI risk & governance
The work of making an AI system aim for what people actually intend, and refuse what they would not accept.
Formal
The field and practice of steering an AI system's goals and behaviour to match the intentions and values of its makers and users - helpful, honest, refusing harmful requests - including in situations it was never tested on.
In plain English
Like King Midas, who wished that everything he touched would turn to gold and got exactly that - including his food. The wish was granted to the letter, not to the meaning.
In practice
A research team at a Danish university adapts an open model to Danish, then has people rate its answers so it learns to admit when it is unsure and to refuse step-by-step help with building weapons.
Why it matters
A capable system that pursues a slightly wrong goal can cause harm at scale; as AI agents act more on their own, the gap between what we asked for and what we meant matters more.