Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Read CSV with Explicit Schema

PySpark data engineering interview problem. Difficulty: beginner. Pattern: Schema Drift. About 12 minutes. Free to practice.

Build a typed events frame from string columns without StructType CSV reads. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

The Lakebench PySpark tab is a pandas-backed shim: there is no real StructType CSV reader. Start from the provided string-typed raw frame (stand-in for a CSV). Cast user_id to int and amount to double. Keep event_type and ts as strings. Order by user_id. Assign the DataFrame to result. Do not call .collect() or .toPandas().

Constraints

  • Fixture has at most 6 rows.
  • Assign result; no collect/toPandas.
  • Cast user_id to int and amount to double.

Examples

Input: raw rows with string user_id/amount Output: user_id | event_type | amount | ts 1 | purchase | 12.5 | 2024-01-01 10:00:00 2 | view | 3.0 | 2024-01-01 11:00:00 3 | purchase | 99.9 | 2024-01-02 09:00:00 4 | refund | 15.0 | 2024-01-02 10:00:00 Cast restores numeric types while preserving event_type and ts.

Topics: lakebench, pyspark, cast, schema.

More PySpark interview questions · All interview problems · Learn data engineering

beginner

Read CSV with Explicit Schema

Interview-style drill: Build a typed events frame from string columns without StructType CSV reads.

The Lakebench PySpark tab is a pandas-backed shim: there is no real `StructType` CSV reader. Start from the provided string-typed `raw` frame (stand-in for a CSV). Cast `user_id` to int and `amount` to double. Keep `event_type` and `ts` as strings. Order by `user_id`. Assign the DataFrame to `result`. Do not call `.collect()` or `.toPandas()`.