PySpark data engineering interview problem. Difficulty: beginner. Pattern: Filtering. About 12 minutes. Free to practice.
Return the top 3 rows by amount descending. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
From df, sort by amount descending and return the top 3 rows (fixture is tiny; production wording would say top 10). Assign result.
Input: events by amount Output: user_id | event_type | amount | ts 2 | purchase | 220.0 | 2024-01-02 12:00:00 1 | purchase | 150.0 | 2024-01-01 10:00:00 3 | purchase | 110.0 | 2024-01-03 09:30:00 Descending amount then limit 3.
Topics: lakebench, pyspark, orderBy, limit.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Return the top 3 rows by amount descending.
From `df`, sort by `amount` descending and return the top 3 rows (fixture is tiny; production wording would say top 10). Assign `result`.