Pandas data engineering interview problem. Difficulty: beginner. Pattern: Missing Data. About 12 minutes. Free to practice.
Production ticket: expect_column_not_null on telemetry.reading - return the actual violating rows, not just a pass/fail. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
A not-null gate on df_telemetry["reading"] is failing. Finance wants the actual rows, not just the boolean. Return the rows of df_telemetry where reading is null. Columns: reading_id, device_id, plant, metric (drop reading, event_time, ingest_time). Sort by reading_id. Assign result.
Input: reading_id | device_id | plant | metric | reading R00008 | DEV-008 | chennai-east | temp_c | (null) R00009 | DEV-009 | pune-west | vibration_mm | 0.412 Output: reading_id | device_id | plant | metric R00008 | DEV-008 | chennai-east | temp_c R00008 has no reading and is kept; R00009 has a real value and is dropped.
Topics: data-quality, not-null, pandas.
More interview problems · All interview problems · Learn data engineering
Production ticket: expect_column_not_null on telemetry.reading - return the actual violating rows, not just a pass/fail.
A not-null gate on `df_telemetry["reading"]` is failing. Finance wants the actual rows, not just the boolean. Return the rows of `df_telemetry` where `reading` is null. Columns: `reading_id`, `device_id`, `plant`, `metric` (drop `reading`, `event_time`, `ingest_time`). Sort by `reading_id`. Assign `result`.