Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Rank vs Dense Rank

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Window Functions. About 16 minutes. Part of the Pro drill bank.

Show rank, dense_rank, and row_number per user by amount desc. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

For each user_id, order by amount desc then ts and add rank_r, dense_r, row_r. Select those plus user_id, amount, ts. Order by user_id, row_r. Assign result.

Constraints

  • Include all three ranking columns.
  • Break amount ties with ts.

Examples

Input: tied amounts for user 1 Output: user_id | amount | ts | rank_r | dense_r | row_r 1 | 40.0 | 2024-01-01 10:00:00 | 1 | 1 | 1 1 | 40.0 | 2024-01-01 11:00:00 | 1 | 1 | 2 1 | 10.0 | 2024-01-01 09:00:00 | 3 | 2 | 3 Each user's eligible order totals are aggregated; users with higher totals appear earlier in the ranked output.

Topics: lakebench, pyspark, rank, dense_rank, row_number.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Rank vs Dense Rank

Interview-style drill: Show rank, dense_rank, and row_number per user by amount desc.

For each `user_id`, order by `amount` desc then `ts` and add `rank_r`, `dense_r`, `row_r`. Select those plus `user_id`, `amount`, `ts`. Order by `user_id`, `row_r`. Assign `result`.