Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Forward-fill readings per device

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Window Functions. About 16 minutes. Part of the Pro drill bank.

Replace NULL readings with the last non-null reading of the same device, in time order. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

readings has device, ts (a timestamp string) and reading, which is sometimes NULL. The rows are not in time order. Add a column filled: the most recent non-null reading of the same device at or before that row's ts. If a device has no earlier non-null reading, filled is NULL. Keep all original columns. Assign the DataFrame to result.

Requirements

  • Fill never crosses devices.
  • Original columns are kept.

Constraints

  • ts strings sort in time order.
  • Each device has unique timestamps.

Examples

Input: readings device | ts | reading d1 | 2024-01-01 09:20:00 | NULL d1 | 2024-01-01 09:00:00 | 10.5 d1 | 2024-01-01 09:30:00 | 11 d1 | 2024-01-01 09:10:00 | NULL d2 | 2024-01-01 09:00:00 | NULL d2 | 2024-01-01 09:10:00 | 7.5 d2 | 2024-01-01 09:20:00 | NULL d3 | 2024-01-01 09:00:00 | NULL Output: device | ts | reading | filled d1 | 2024-01-01 09:00:00 | 10.5 | 10.5 d1 | 2024-01-01 09:10:00 | NULL | 10.5 d1 | 2024-01-01 09:20:00 | NULL | 10.5 d1 | 2024-01-01 09:30:00 | 11 | 11 d2 | 2024-01-01 09:00:00 | NULL | NULL d2 | 2024-01-01 09:10:00 | 7.5 | 7.5 d2 | 2024-01-01 09:20:00 | NULL | 7.5 d3 | 2024-01-01 09:00:00 | NULL | NULL d1 at 09:10 and 09:20 take 10.5, d2 at 09:20 takes 7.5, and the leading NULL rows (d2 at 09:00 and all of d3) stay NULL.

Topics: lakebench, pyspark, last, ignorenulls, window.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Forward-fill readings per device

Interview-style drill: Replace NULL readings with the last non-null reading of the same device, in time order.

`readings` has `device`, `ts` (a timestamp string) and `reading`, which is sometimes NULL. The rows are not in time order. Add a column `filled`: the most recent non-null `reading` of the **same device** at or before that row's `ts`. If a device has no earlier non-null reading, `filled` is NULL. Keep all original columns. Assign the DataFrame to `result`.