PySpark data engineering interview problem. Difficulty: advanced. Pattern: Medallion. About 20 minutes. Part of the Pro drill bank.
Repartition by date and region instead of writing Parquet partitions. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
There is no Parquet writer here. repartition(3, 'date', 'region') to mimic partition columns, then select all fields ordered by id. Assign result.
Input: sales by date/region Output: id | date | region | amount 1 | 2024-01-01 | US | 10.0 2 | 2024-01-01 | EU | 20.0 3 | 2024-01-02 | US | 30.0 4 | 2024-01-02 | EU | 40.0 5 | 2024-01-03 | US | 50.0 6 | 2024-01-03 | APAC | 60.0 Shim prints a repartition note; row set stays intact.
Topics: lakebench, pyspark, repartition.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Repartition by date and region instead of writing Parquet partitions.
There is no Parquet writer here. `repartition(3, 'date', 'region')` to mimic partition columns, then select all fields ordered by `id`. Assign `result`.