Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Infer Schema vs Explicit Schema

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Schema Drift. About 16 minutes. Part of the Pro drill bank.

Cast string columns and drop failed numeric casts. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

No StructType reader here. Cast user_id to int and amount to double from raw, then keep rows where amount is not null (bad casts become null). Order by user_id. Assign result.

Constraints

  • No StructType API.

Examples

Input: string raw Output: user_id | amount 1 | 10.5 2 | 20.0 4 | 40.25 Row 3 casts to null and drops.

Topics: lakebench, pyspark, cast, schema.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Infer Schema vs Explicit Schema

Interview-style drill: Cast string columns and drop failed numeric casts.

No StructType reader here. Cast `user_id` to int and `amount` to double from `raw`, then keep rows where `amount` is not null (bad casts become null). Order by `user_id`. Assign `result`.