PySpark data engineering interview problem. Difficulty: beginner. Pattern: Filtering. About 12 minutes. Free to practice.
Explode tags arrays into one row per tag. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
From df, explode tags into a tag column. Return user_id, tag ordered by both. Assign result.
Input: user_id | tags Output: user_id | tag 1 | spark 1 | sql 2 | pandas 3 | flink 3 | kafka 3 | spark 4 | dbt Array grain becomes tag grain; other columns repeat.
Topics: lakebench, pyspark, explode.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Explode tags arrays into one row per tag.
From `df`, explode `tags` into a `tag` column. Return `user_id`, `tag` ordered by both. Assign `result`.