Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Forward-fill per device

Pandas data engineering interview problem. Difficulty: intermediate. Pattern: Missing Data. About 14 minutes. Part of the Pro drill bank.

Fill missing readings from the same device only, never across devices. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

df has device, ts and reading; reading is sometimes missing and the rows are not in time order. Fill each missing reading with the most recent earlier reading of the same device. A device with no earlier reading keeps the missing value. Keep the three columns, sorted by device then ts, with a clean index. Assign the DataFrame to result.

Requirements

  • Never fill across devices.

Constraints

  • ts strings sort in time order.
  • Timestamps are unique per device.

Examples

Input: df device | ts | reading d1 | 09:20 | NULL d1 | 09:00 | 10.5 d1 | 09:10 | NULL d2 | 09:10 | 7.5 d2 | 09:00 | NULL d2 | 09:20 | NULL d3 | 09:00 | NULL Output: device | ts | reading d1 | 09:00 | 10.5 d1 | 09:10 | 10.5 d1 | 09:20 | 10.5 d2 | 09:00 | NULL d2 | 09:10 | 7.5 d2 | 09:20 | 7.5 d3 | 09:00 | NULL d1's gaps take 10.5. d2's first reading is missing and has nothing earlier; its later gap takes 7.5. d3 only has a missing value.

Topics: lakebench, pandas, ffill, groupby, missing data.

More interview problems · All interview problems · Learn data engineering

intermediate

Forward-fill per device

Interview-style drill: Fill missing readings from the same device only, never across devices.

`df` has `device`, `ts` and `reading`; `reading` is sometimes missing and the rows are not in time order. Fill each missing `reading` with the most recent earlier reading of the **same device**. A device with no earlier reading keeps the missing value. Keep the three columns, sorted by `device` then `ts`, with a clean index. Assign the DataFrame to `result`.