Python data engineering interview problem. Difficulty: intermediate. Pattern: Deduplication. About 15 minutes. Part of the Pro drill bank.
A CRM export repeats the same customer under slightly different keys. Dedup must collapse those rows using the fuzzy key the prompt defines, not a naive exact match only.
Deduplicate customer records by the fuzzy key described in the prompt.
Input: dedupe_customers([{'name': ' Ann ', 'zip': '94110', 'updated_at': '2026-01-01', 'id': 1}, {'name': 'ann', 'zip': '94110', 'updated_at': '2026-02-01', 'id': 2}]) Output: [{'name': 'ann', 'zip': '94110', 'updated_at': '2026-02-01', 'id': 2}] This input follows the stated rules and produces this output.
Topics: lakebench, python, deduplication.
More Python interview questions · All interview problems · Learn data engineering
Interview-style drill: Given customer records with a normalized name+zip key, keep only the most recently updated record per key.
Deduplicate customer records by the fuzzy key described in the prompt.