Define the evaluation question
Course overview · 4 min reading + 12 min practice, estimated
Principles and method
Evaluate a specific task and use context, not whether a model is generally good. State the inputs, expected outputs, users and consequences of error. Define unacceptable failures and acceptable variation. A summary can use different wording while preserving facts; a fabricated qualification is materially different. Include omissions, unsupported inferences and privacy failures in the rubric. Establish the acceptance decision before seeing results. Supplier benchmarks can inform questions but do not replace evidence from your workflow and representative inputs.
Worked example
The evaluation asks whether an assistant produces evidence-linked interview packs from approved records without adding claims or mixing people. It does not ask reviewers whether the writing sounds impressive.
Put it into practice
Write an evaluation charter with three quality dimensions and two blocking failures.
Use fictional information and keep your work in your own notes.
Compare your approach: self-review guidance
Tie each dimension to the task’s purpose. A blocking failure should have a clear consequence and stop rule, not merely a low average score.
Sources and further reading
Original Academy teaching and fictional examples. These references provide context, not endorsement. Edition 2026.09; updated 2026-09-24.
- GOV.UK: Responsible AI in recruitment
UK guidance on procuring and deploying recruitment AI.
- NIST: AI Risk Management Framework
Voluntary framework for organising AI risks and controls.