Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Partition by Column

PySpark data engineering interview problem. Difficulty: advanced. Pattern: Medallion. About 20 minutes. Part of the Pro drill bank.

Repartition by date and region instead of writing Parquet partitions. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

There is no Parquet writer here. repartition(3, 'date', 'region') to mimic partition columns, then select all fields ordered by id. Assign result.

Constraints

  • Pass date and region to repartition.
  • Keep all columns.

Examples

Input: sales by date/region Output: id | date | region | amount 1 | 2024-01-01 | US | 10.0 2 | 2024-01-01 | EU | 20.0 3 | 2024-01-02 | US | 30.0 4 | 2024-01-02 | EU | 40.0 5 | 2024-01-03 | US | 50.0 6 | 2024-01-03 | APAC | 60.0 Shim prints a repartition note; row set stays intact.

Topics: lakebench, pyspark, repartition.

More PySpark interview questions · All interview problems · Learn data engineering

advanced

Partition by Column

Interview-style drill: Repartition by date and region instead of writing Parquet partitions.

There is no Parquet writer here. `repartition(3, 'date', 'region')` to mimic partition columns, then select all fields ordered by `id`. Assign `result`.