PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Schema Drift. About 16 minutes. Part of the Pro drill bank.
Cast string columns and drop failed numeric casts. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
No StructType reader here. Cast user_id to int and amount to double from raw, then keep rows where amount is not null (bad casts become null). Order by user_id. Assign result.
Input: string raw Output: user_id | amount 1 | 10.5 2 | 20.0 4 | 40.25 Row 3 casts to null and drops.
Topics: lakebench, pyspark, cast, schema.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Cast string columns and drop failed numeric casts.
No StructType reader here. Cast `user_id` to int and `amount` to double from `raw`, then keep rows where `amount` is not null (bad casts become null). Order by `user_id`. Assign `result`.