Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Validate Data Types

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Data Quality. About 16 minutes. Part of the Pro drill bank.

Return mismatched columns against an expected type map (empty if clean). Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Validate that user_id behaves as int and amount as double for this fixture. Return columns column_name, expected_type, actual_type for mismatches only (empty frame if none). Assign result.

Constraints

  • Mismatch columns only when types disagree.

Examples

Input: typed df Output: column_name | expected_type | actual_type No mismatches for the provided fixture.

Topics: lakebench, pyspark, schema, validation.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Validate Data Types

Interview-style drill: Return mismatched columns against an expected type map (empty if clean).

Validate that `user_id` behaves as int and `amount` as double for this fixture. Return columns `column_name`, `expected_type`, `actual_type` for mismatches only (empty frame if none). Assign `result`.