Pandas data engineering interview problem. Difficulty: beginner. Pattern: Deduplication. About 10 minutes. Part of the Pro drill bank.
Keep only the most recent record of each customer. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
df holds several versions of each customer: customer_id, name, email and updated_at (an ISO date string). Return one row per customer_id: the version with the latest updated_at. Keep all four columns, sorted by customer_id, with a clean 0..n-1 index. Assign the DataFrame to result.
Input: df customer_id | name | email | updated_at 1 | Ann | ann@v3.com | 2024-03-01 2 | Bo | bo@new.com | 2024-03-10 1 | Ann | ann@v1.com | 2024-01-01 3 | Cy | cy@x.com | 2024-01-15 1 | Ann | ann@v2.com | 2024-02-01 2 | Bo | bo@old.com | 2024-02-01 Output: customer_id | name | email | updated_at 1 | Ann | ann@v3.com | 2024-03-01 2 | Bo | bo@new.com | 2024-03-10 3 | Cy | cy@x.com | 2024-01-15 Customer 1's newest version is the March one even though it arrives first; customer 2's newest is 10 March; customer 3 has only one row.
Topics: lakebench, pandas, drop_duplicates, latest, sort.
More interview problems · All interview problems · Learn data engineering
Interview-style drill: Keep only the most recent record of each customer.
`df` holds several versions of each customer: `customer_id`, `name`, `email` and `updated_at` (an ISO date string). Return one row per `customer_id`: the version with the latest `updated_at`. Keep all four columns, sorted by `customer_id`, with a clean 0..n-1 index. Assign the DataFrame to `result`.