PySpark data engineering interview problem. Difficulty: advanced. Pattern: Skew. About 20 minutes. Part of the Pro drill bank.
Detect hot keys via per-user row counts (no Spark UI). Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Group by user_id and count rows as row_count. Order by row_count desc, then user_id. The hot key is obvious from counts. Assign result.
Input: skewed keys Output: user_id | row_count 1 | 4 2 | 1 3 | 1 user_id 1 dominates the fixture.
Topics: lakebench, pyspark, skew, counts.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Detect hot keys via per-user row counts (no Spark UI).
Group by `user_id` and count rows as `row_count`. Order by `row_count` desc, then `user_id`. The hot key is obvious from counts. Assign `result`.