PySpark data engineering interview problem. Difficulty: beginner. Pattern: Schema Drift. About 12 minutes. Free to practice.
Build a typed events frame from string columns without StructType CSV reads. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
The Lakebench PySpark tab is a pandas-backed shim: there is no real StructType CSV reader. Start from the provided string-typed raw frame (stand-in for a CSV). Cast user_id to int and amount to double. Keep event_type and ts as strings. Order by user_id. Assign the DataFrame to result. Do not call .collect() or .toPandas().
Input: raw rows with string user_id/amount Output: user_id | event_type | amount | ts 1 | purchase | 12.5 | 2024-01-01 10:00:00 2 | view | 3.0 | 2024-01-01 11:00:00 3 | purchase | 99.9 | 2024-01-02 09:00:00 4 | refund | 15.0 | 2024-01-02 10:00:00 Cast restores numeric types while preserving event_type and ts.
Topics: lakebench, pyspark, cast, schema.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Build a typed events frame from string columns without StructType CSV reads.
The Lakebench PySpark tab is a pandas-backed shim: there is no real `StructType` CSV reader. Start from the provided string-typed `raw` frame (stand-in for a CSV). Cast `user_id` to int and `amount` to double. Keep `event_type` and `ts` as strings. Order by `user_id`. Assign the DataFrame to `result`. Do not call `.collect()` or `.toPandas()`.