PySpark data engineering interview problem. Difficulty: advanced. Pattern: Streaming. About 20 minutes. Part of the Pro drill bank.
Parse Kafka-like JSON payloads into columns (no Kafka source). Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
From raw.value JSON strings, extract user_id (int), event_type, ts with get_json_object. Order by ts. Assign result.
Input: json values Output: user_id | event_type | ts 1 | click | 2024-01-01 10:00:00 2 | view | 2024-01-01 10:01:00 1 | purchase | 2024-01-01 10:02:00 3 | click | 2024-01-01 10:03:00 JSON fields become columns.
Topics: lakebench, pyspark, json, get_json_object.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Parse Kafka-like JSON payloads into columns (no Kafka source).
From `raw.value` JSON strings, extract `user_id` (int), `event_type`, `ts` with `get_json_object`. Order by `ts`. Assign `result`.