Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Pandas UDF (Vectorized)

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Filtering. About 16 minutes. Part of the Pro drill bank.

Same banding as problem 33; pandas_udf is unsupported in the shim. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Rewrite the banding without pandas_udf (unsupported here). Use the same when rules as problem 33. Return user_id, amount, band ordered by user_id. Assign result.

Constraints

  • Do not call pandas_udf.
  • Labels are Low/Medium/High.

Examples

Input: amounts Output: user_id | amount | band 1 | 20.0 | Low 4 | 250.0 | High Native expressions avoid Python row workers.

Topics: lakebench, pyspark, when, pandas_udf-reframe.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Pandas UDF (Vectorized)

Interview-style drill: Same banding as problem 33; pandas_udf is unsupported in the shim.

Rewrite the banding without `pandas_udf` (unsupported here). Use the same `when` rules as problem 33. Return `user_id`, `amount`, `band` ordered by `user_id`. Assign `result`.