Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Null count for every column

PySpark data engineering interview problem. Difficulty: beginner. Pattern: Data Quality. About 10 minutes. Part of the Pro drill bank.

Build a one-row profile of NULL counts for all columns, without naming the columns. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

df has several columns, some of them with NULL values. Return a one-row DataFrame that has the same column names as df, where each value is the number of NULLs in that column. Build it from df.columns so the code keeps working when columns are added. Assign the DataFrame to result.

Requirements

  • One output row.
  • Same column names as the input.

Constraints

  • Column names are not known in advance.
  • Counts are integers.

Examples

Input: df id | name | score 1 | a | 10 2 | NULL | NULL 3 | c | NULL NULL | NULL | 5 Output: id | name | score 1 | 2 | 2 id is NULL once, name twice and score twice.

Topics: lakebench, pyspark, nulls, profiling, columns.

More PySpark interview questions · All interview problems · Learn data engineering

beginner

Null count for every column

Interview-style drill: Build a one-row profile of NULL counts for all columns, without naming the columns.

`df` has several columns, some of them with NULL values. Return a one-row DataFrame that has the same column names as `df`, where each value is the number of NULLs in that column. Build it from `df.columns` so the code keeps working when columns are added. Assign the DataFrame to `result`.