PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Window Functions. About 16 minutes. Part of the Pro drill bank.
Replace NULL readings with the last non-null reading of the same device, in time order. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
readings has device, ts (a timestamp string) and reading, which is sometimes NULL. The rows are not in time order. Add a column filled: the most recent non-null reading of the same device at or before that row's ts. If a device has no earlier non-null reading, filled is NULL. Keep all original columns. Assign the DataFrame to result.
Input: readings device | ts | reading d1 | 2024-01-01 09:20:00 | NULL d1 | 2024-01-01 09:00:00 | 10.5 d1 | 2024-01-01 09:30:00 | 11 d1 | 2024-01-01 09:10:00 | NULL d2 | 2024-01-01 09:00:00 | NULL d2 | 2024-01-01 09:10:00 | 7.5 d2 | 2024-01-01 09:20:00 | NULL d3 | 2024-01-01 09:00:00 | NULL Output: device | ts | reading | filled d1 | 2024-01-01 09:00:00 | 10.5 | 10.5 d1 | 2024-01-01 09:10:00 | NULL | 10.5 d1 | 2024-01-01 09:20:00 | NULL | 10.5 d1 | 2024-01-01 09:30:00 | 11 | 11 d2 | 2024-01-01 09:00:00 | NULL | NULL d2 | 2024-01-01 09:10:00 | 7.5 | 7.5 d2 | 2024-01-01 09:20:00 | NULL | 7.5 d3 | 2024-01-01 09:00:00 | NULL | NULL d1 at 09:10 and 09:20 take 10.5, d2 at 09:20 takes 7.5, and the leading NULL rows (d2 at 09:00 and all of d3) stay NULL.
Topics: lakebench, pyspark, last, ignorenulls, window.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Replace NULL readings with the last non-null reading of the same device, in time order.
`readings` has `device`, `ts` (a timestamp string) and `reading`, which is sometimes NULL. The rows are not in time order. Add a column `filled`: the most recent non-null `reading` of the **same device** at or before that row's `ts`. If a device has no earlier non-null reading, `filled` is NULL. Keep all original columns. Assign the DataFrame to `result`.