AI Output Evaluation and Quality Assurance · Lesson 2

Build a representative and challenging case set

Course overview · 4 min reading + 12 min practice, estimated

Principles and method

Use fictional or appropriately governed data that reflects the task’s real variation. Include ordinary records, sparse records, conflicting dates, unusual formats and deliberate untrusted instructions. Separate development cases from a held-out evaluation set so repeated prompt tuning does not simply memorise the test. Document what the set covers and what it misses. Rare but serious failures may need targeted challenge cases, reported separately from representative performance. Do not claim a population error rate from a handpicked adversarial set.

Worked example

A test set includes multilingual job titles, two people with similar names and a CV containing instructions to reveal other records. The report distinguishes routine cases from security challenges.

Put it into practice

Specify twelve test cases and split them into development, held-out and challenge groups.

Use fictional information and keep your work in your own notes.

Compare your approach: self-review guidance

Explain the coverage purpose of each case. Include at least one omission risk and one cross-record confusion risk. Keep the final test cases unchanged while selecting a prompt.

Download the course workbook

Sources and further reading

Original Academy teaching and fictional examples. These references provide context, not endorsement. Edition 2026.09; updated 2026-09-24.

How our learning is designed