PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Filtering. About 16 minutes. Part of the Pro drill bank.
Band amounts Low/Medium/High with when (no Python UDF). Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Interview prompt often asks for a UDF. Here, band amount as Low (<50), Medium (50–200), High (>200) using when/otherwise only. Return user_id, amount, band ordered by user_id. Assign result.
Input: amounts Output: user_id | amount | band 1 | 20.0 | Low 2 | 80.0 | Medium 3 | 150.0 | Medium 4 | 250.0 | High 5 | 50.0 | Medium 6 | 200.0 | Medium Inclusive Medium on both endpoints.
Topics: lakebench, pyspark, when, udf-reframe.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Band amounts Low/Medium/High with when (no Python UDF).
Interview prompt often asks for a UDF. Here, band `amount` as Low (<50), Medium (50–200), High (>200) using `when`/`otherwise` only. Return `user_id`, `amount`, `band` ordered by `user_id`. Assign `result`.