Evaluation & metrics
The practice of checking how well an AI system does its job, with scores, public tests, human review and deliberate attacks.
Formal
The planned measuring of a model or AI system against goals for quality, safety and fairness, combining numbers such as accuracy and the F1 score, benchmark results, human review, LLM-as-a-judge grading and AI red teaming.
In plain English
Like a car's road trials before sale, with speed runs on the track, crash checks, drivers' opinions and people trying hard to make it fail.
In practice
Before moving its tenant chat to a newer large language model, a housing association's IT lead reruns 500 saved tenant questions, has staff grade a sample of the answers, and switches only if nothing gets worse.
Why it matters
Without regular checks nobody knows whether a change made the system better, worse or unsafe, and the EU AI Act demands documented testing for systems used in areas such as hiring, credit scoring or welfare decisions.