Data Quality, Taxonomies and Record Hygiene · Lesson 3

Resolve duplicates without losing meaning

Course overview · 4 min reading + 12 min practice, estimated

Principles and method

Duplicate detection should generate candidates for review, not assume identity from a name. Use appropriate combinations of identifiers and context while respecting data minimisation. A merge needs rules for conflicting fields, provenance, application history and contact preferences. The most recent value is not always the most reliable. Preserve restrictions so a suppressed record cannot become contactable through a merge. Keep a reversible or auditable process where practical and test on fictional collisions before bulk changes. Never silently combine records that may belong to different people.

Worked example

Two records share a name but have different employment histories. The reviewer keeps them separate pending verification. Another confirmed duplicate retains the stricter contact preference and both source histories after merging.

Put it into practice

Write a merge decision tree for three fictional pairs, including one uncertain identity.

Use fictional information and keep your work in your own notes.

Compare your approach: self-review guidance

Include evidence needed, who can approve, field conflict rules and downstream checks. Uncertain identity should remain unresolved rather than being forced into a clean-looking database.

Download the course workbook

Sources and further reading

Original Academy teaching and fictional examples. These references provide context, not endorsement. Edition 2026.09; updated 2026-09-24.

How our learning is designed