Pandas data engineering interview problem. Difficulty: intermediate. Pattern: Missing Data. About 14 minutes. Part of the Pro drill bank.
Fill missing readings from the same device only, never across devices. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
df has device, ts and reading; reading is sometimes missing and the rows are not in time order. Fill each missing reading with the most recent earlier reading of the same device. A device with no earlier reading keeps the missing value. Keep the three columns, sorted by device then ts, with a clean index. Assign the DataFrame to result.
Input: df device | ts | reading d1 | 09:20 | NULL d1 | 09:00 | 10.5 d1 | 09:10 | NULL d2 | 09:10 | 7.5 d2 | 09:00 | NULL d2 | 09:20 | NULL d3 | 09:00 | NULL Output: device | ts | reading d1 | 09:00 | 10.5 d1 | 09:10 | 10.5 d1 | 09:20 | 10.5 d2 | 09:00 | NULL d2 | 09:10 | 7.5 d2 | 09:20 | 7.5 d3 | 09:00 | NULL d1's gaps take 10.5. d2's first reading is missing and has nothing earlier; its later gap takes 7.5. d3 only has a missing value.
Topics: lakebench, pandas, ffill, groupby, missing data.
More interview problems · All interview problems · Learn data engineering
Interview-style drill: Fill missing readings from the same device only, never across devices.
`df` has `device`, `ts` and `reading`; `reading` is sometimes missing and the rows are not in time order. Fill each missing `reading` with the most recent earlier reading of the **same device**. A device with no earlier reading keeps the missing value. Keep the three columns, sorted by `device` then `ts`, with a clean index. Assign the DataFrame to `result`.