# AI Output Evaluation and Quality Assurance: practice workbook
Talent Engineering Academy | An education initiative by Vitae
Edition 2026.09 | Updated 2026-09-24
Course: https://talentengineering.org/courses/ai-evaluation

Build representative test sets, score material errors and decide whether an AI workflow is ready for its stated use.

## Your deliverable
An evaluation dataset specification, rubric and release decision.

Use fictional information. Keep your completed work in your own secure notes. Exercises and capstone work are self-directed, not independently assessed.

## 1. Intended use and failure consequences

Your notes:



## 2. Representative and challenge cases

Your notes:



## 3. Reference evidence and rubric

Your notes:



## 4. Independent review and adjudication

Your notes:



## 5. Results, uncertainty and acceptance gates

Your notes:



## 6. Regression set and change triggers

Your notes:



## Lesson exercises

### 1. Define the evaluation question

Write an evaluation charter with three quality dimensions and two blocking failures.

Your response:


Worked example: The evaluation asks whether an assistant produces evidence-linked interview packs from approved records without adding claims or mixing people. It does not ask reviewers whether the writing sounds impressive.

Self-review guidance: Tie each dimension to the task’s purpose. A blocking failure should have a clear consequence and stop rule, not merely a low average score.

### 2. Build a representative and challenging case set

Specify twelve test cases and split them into development, held-out and challenge groups.

Your response:


Worked example: A test set includes multilingual job titles, two people with similar names and a CV containing instructions to reveal other records. The report distinguishes routine cases from security challenges.

Self-review guidance: Explain the coverage purpose of each case. Include at least one omission risk and one cross-record confusion risk. Keep the final test cases unchanged while selecting a prompt.

### 3. Score outputs with an auditable rubric

Score five fictional outputs using a four-category rubric and write an adjudication note.

Your response:


Worked example: Nineteen outputs are accurate and one includes another candidate’s contact details. A 95% pass rate is not enough to approve a workflow whose gate prohibits cross-record disclosure.

Self-review guidance: Record material errors individually. Explain why one serious failure can block release even when the average looks good. Preserve the source references needed to reproduce the judgement.

### 4. Make a release and regression decision

Write a release decision for mixed fictional results and a regression schedule.

Your response:


Worked example: A new model improves readability but fails two evidence-reference checks. The team retains the earlier approved version for the task while investigating, rather than upgrading solely because the model is newer.

Self-review guidance: Reference the gates, unresolved risks and permitted scope. Define what change triggers retesting and who can stop use. Do not alter the gates after seeing the result merely to obtain a pass.

## Portfolio review
Check that your work is internally consistent, distinguishes facts from assumptions, names decision owners and explains its limitations. Revise gaps before using the method in real work.

## Further reading
- [GOV.UK: Responsible AI in recruitment](https://www.gov.uk/government/publications/responsible-ai-in-recruitment-guide/responsible-ai-in-recruitment): UK guidance on procuring and deploying recruitment AI.
- [NIST: AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework): Voluntary framework for organising AI risks and controls.

Original educational scenarios. References provide further reading and do not imply endorsement. Check current official rules and appropriate professional advice for real legal, financial or regulated decisions.