PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Filtering. About 16 minutes. Part of the Pro drill bank.
Prefer when/otherwise over UDFs for simple banding. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Band amounts with built-in when (same Low/Medium/High rules). Explain in the hint why this beats a Python UDF on a cluster, but grade the DataFrame. Assign result ordered by user_id.
Input: amounts Output: user_id | amount | band 1 | 20.0 | Low 2 | 80.0 | Medium 4 | 250.0 | High Built-ins avoid per-row Python.
Topics: lakebench, pyspark, when.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Prefer when/otherwise over UDFs for simple banding.
Band amounts with built-in `when` (same Low/Medium/High rules). Explain in the hint why this beats a Python UDF on a cluster, but grade the DataFrame. Assign `result` ordered by `user_id`.