Resolve duplicates without losing meaning
Course overview · 4 min reading + 12 min practice, estimated
Principles and method
Duplicate detection should generate candidates for review, not assume identity from a name. Use appropriate combinations of identifiers and context while respecting data minimisation. A merge needs rules for conflicting fields, provenance, application history and contact preferences. The most recent value is not always the most reliable. Preserve restrictions so a suppressed record cannot become contactable through a merge. Keep a reversible or auditable process where practical and test on fictional collisions before bulk changes. Never silently combine records that may belong to different people.
Worked example
Two records share a name but have different employment histories. The reviewer keeps them separate pending verification. Another confirmed duplicate retains the stricter contact preference and both source histories after merging.
Put it into practice
Write a merge decision tree for three fictional pairs, including one uncertain identity.
Use fictional information and keep your work in your own notes.
Compare your approach: self-review guidance
Include evidence needed, who can approve, field conflict rules and downstream checks. Uncertain identity should remain unresolved rather than being forced into a clean-looking database.
Sources and further reading
Original Academy teaching and fictional examples. These references provide context, not endorsement. Edition 2026.09; updated 2026-09-24.
- GOV.UK: Responsible AI in recruitment
UK guidance on procuring and deploying recruitment AI.
- NIST: AI Risk Management Framework
Voluntary framework for organising AI risks and controls.