# Data Quality, Taxonomies and Record Hygiene: practice workbook
Talent Engineering Academy | An education initiative by Vitae
Edition 2026.09 | Updated 2026-09-24
Course: https://talentengineering.org/courses/data-quality-taxonomies

Build reliable field definitions, resolve duplicates carefully and maintain useful skills and role vocabularies.

## Your deliverable
A data-quality rule set and taxonomy maintenance plan.

Use fictional information. Keep your completed work in your own secure notes. Exercises and capstone work are self-directed, not independently assessed.

## 1. Critical fields and quality dimensions

Your notes:



## 2. Vocabulary terms, synonyms and scope

Your notes:



## 3. Duplicate evidence and merge authority

Your notes:



## 4. Validation and exception queue

Your notes:



## 5. Freshness and correction rules

Your notes:



## 6. Quality measures and maintenance owner

Your notes:



## Lesson exercises

### 1. Prioritise quality by decision impact

Rank ten fictional fields by consequence and define checks for the top four.

Your response:


Worked example: A vacancy has a complete salary field but the value is monthly while reports assume annual. The quality rule checks units and meaning, not just whether a cell is filled.

Self-review guidance: Include semantic checks such as currency and time basis. Explain when each value can reasonably be known and how missing information is routed for follow-up.

### 2. Build taxonomies that support retrieval

Create a ten-term vocabulary with synonyms and two examples of terms that must remain distinct.

Your response:


Worked example: PLC programming and automation engineering overlap but are not interchangeable. The vocabulary links related concepts while recording specific evidence and proficiency context separately.

Self-review guidance: Define inclusion examples and ambiguous cases. Avoid treating a keyword match as verified competence. Show how an unfamiliar term can be reviewed without immediately creating a duplicate label.

### 3. Resolve duplicates without losing meaning

Write a merge decision tree for three fictional pairs, including one uncertain identity.

Your response:


Worked example: Two records share a name but have different employment histories. The reviewer keeps them separate pending verification. Another confirmed duplicate retains the stricter contact preference and both source histories after merging.

Self-review guidance: Include evidence needed, who can approve, field conflict rules and downstream checks. Uncertain identity should remain unresolved rather than being forced into a clean-looking database.

### 4. Make hygiene a maintained workflow

Design a weekly hygiene routine and an import failure response.

Your response:


Worked example: An import repeatedly converts unknown availability into immediately available. The team stops the import, fixes the mapping and reviews affected records instead of asking recruiters to correct them indefinitely by hand.

Self-review guidance: Include prevention, exception ownership and validation after repair. A useful routine focuses on consequential defects and does not reward staff for making uncertain fields look complete.

## Portfolio review
Check that your work is internally consistent, distinguishes facts from assumptions, names decision owners and explains its limitations. Revise gaps before using the method in real work.

## Further reading
- [GOV.UK: Responsible AI in recruitment](https://www.gov.uk/government/publications/responsible-ai-in-recruitment-guide/responsible-ai-in-recruitment): UK guidance on procuring and deploying recruitment AI.
- [NIST: AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework): Voluntary framework for organising AI risks and controls.

Original educational scenarios. References provide further reading and do not imply endorsement. Check current official rules and appropriate professional advice for real legal, financial or regulated decisions.