# Monitoring, Incident Handling and Continuous Improvement: practice workbook
Talent Engineering Academy | An education initiative by Vitae
Edition 2026.09 | Updated 2026-09-24
Course: https://talentengineering.org/courses/monitoring-improvement

Detect meaningful failures, contain incidents and improve recruitment systems using evidence.

## Your deliverable
An operational dashboard, incident runbook and improvement backlog.

Use fictional information. Keep your completed work in your own secure notes. Exercises and capstone work are self-directed, not independently assessed.

## 1. Service signals and alert thresholds

Your notes:



## 2. Severity and affected people

Your notes:



## 3. Containment and communication owner

Your notes:



## 4. Evidence preservation and investigation

Your notes:



## 5. Recovery and validation

Your notes:



## 6. Root cause, actions and effectiveness review

Your notes:



## Lesson exercises

### 1. Monitor the service and its safeguards

Choose six monitoring signals for an interview-pack workflow and assign actions.

Your response:


Worked example: A workflow reports 99% technical success but reviewers correct many role locations. The dashboard adds a material correction measure and a source-field check rather than celebrating successful requests alone.

Self-review guidance: Include one quality signal, one backlog signal and one handling control. Explain the denominator and threshold. Remove alerts that have no clear owner or response.

### 2. Contain incidents before optimising diagnosis

Write the first thirty minutes of an incident response for a misdirected pack.

Your response:


Worked example: A pack contains another candidate’s contact details. The owner pauses further sends, identifies recipients and affected records, involves the privacy lead and follows the incident process before resuming.

Self-review guidance: Name containment, investigation and communication owners. Keep factual timestamps and avoid speculative assurances. Include how essential candidate coordination continues safely during the pause.

### 3. Recover with validation and clear communication

Create a recovery checklist and a concise internal status update using fictional facts.

Your response:


Worked example: After fixing an attachment mapping bug, the team tests similar names, multiple applications and retries. It reconciles queued work before resuming a limited batch with increased review.

Self-review guidance: Separate verified facts from investigation questions. Include validation evidence, queue reconciliation and a resumption decision. A service restart without these checks is incomplete recovery.

### 4. Learn without stopping at human error

Write a short incident review with two contributing factors and three corrective actions.

Your response:


Worked example: A reviewer approved a wrong attachment because the interface displayed only a filename shared by many records. The action adds identity context and binding checks rather than only telling reviewers to pay more attention.

Self-review guidance: For each action, state the mechanism it changes and the evidence that will show improvement. Include one system design change and a follow-up review date. Avoid blame as a substitute for analysis.

## Portfolio review
Check that your work is internally consistent, distinguishes facts from assumptions, names decision owners and explains its limitations. Revise gaps before using the method in real work.

## Further reading
- [GOV.UK: Responsible AI in recruitment](https://www.gov.uk/government/publications/responsible-ai-in-recruitment-guide/responsible-ai-in-recruitment): UK guidance on procuring and deploying recruitment AI.
- [NIST: AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework): Voluntary framework for organising AI risks and controls.

Original educational scenarios. References provide further reading and do not imply endorsement. Check current official rules and appropriate professional advice for real legal, financial or regulated decisions.