Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Identify Skew

PySpark data engineering interview problem. Difficulty: advanced. Pattern: Skew. About 20 minutes. Part of the Pro drill bank.

Detect hot keys via per-user row counts (no Spark UI). Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Group by user_id and count rows as row_count. Order by row_count desc, then user_id. The hot key is obvious from counts. Assign result.

Constraints

  • Do not invent Spark UI steps.
  • Order by row_count desc.

Examples

Input: skewed keys Output: user_id | row_count 1 | 4 2 | 1 3 | 1 user_id 1 dominates the fixture.

Topics: lakebench, pyspark, skew, counts.

More PySpark interview questions · All interview problems · Learn data engineering

advanced

Identify Skew

Interview-style drill: Detect hot keys via per-user row counts (no Spark UI).

Group by `user_id` and count rows as `row_count`. Order by `row_count` desc, then `user_id`. The hot key is obvious from counts. Assign `result`.