PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Filtering. About 16 minutes. Part of the Pro drill bank.
Same banding as problem 33; pandas_udf is unsupported in the shim. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Rewrite the banding without pandas_udf (unsupported here). Use the same when rules as problem 33. Return user_id, amount, band ordered by user_id. Assign result.
Input: amounts Output: user_id | amount | band 1 | 20.0 | Low 4 | 250.0 | High Native expressions avoid Python row workers.
Topics: lakebench, pyspark, when, pandas_udf-reframe.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Same banding as problem 33; pandas_udf is unsupported in the shim.
Rewrite the banding without `pandas_udf` (unsupported here). Use the same `when` rules as problem 33. Return `user_id`, `amount`, `band` ordered by `user_id`. Assign `result`.