AI Output Evaluation and Quality Assurance · Lesson 3

Score outputs with an auditable rubric

Course overview · 4 min reading + 12 min practice, estimated

Principles and method

Use claim-level checks against reference evidence where factual accuracy matters. Record severity, error type and affected criterion. Separate style preferences from factual defects. Ask reviewers to score independently on a subset and resolve disagreements with the source material. Automated evaluators can help triage, but they can also make mistakes or share a model’s weaknesses. Validate their agreement with appropriate human review before relying on them. Report counts and examples alongside aggregate scores so a good average cannot hide a catastrophic case.

Worked example

Nineteen outputs are accurate and one includes another candidate’s contact details. A 95% pass rate is not enough to approve a workflow whose gate prohibits cross-record disclosure.

Put it into practice

Score five fictional outputs using a four-category rubric and write an adjudication note.

Use fictional information and keep your work in your own notes.

Compare your approach: self-review guidance

Record material errors individually. Explain why one serious failure can block release even when the average looks good. Preserve the source references needed to reproduce the judgement.

Download the course workbook

Sources and further reading

Original Academy teaching and fictional examples. These references provide context, not endorsement. Edition 2026.09; updated 2026-09-24.

How our learning is designed