Pandas data engineering interview problem. Difficulty: beginner. Pattern: Missing Data. About 12 minutes. Free to practice.
Show dropna, mean fill, and ffill strategies labeled by method. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
From df, build three frames: (1) dropna rows, (2) fill score with column mean, (3) forward-fill score. Concat with a method column in {drop,mean,ffill}. Sort by method, user_id. Assign result.
Input: scores with nulls Output: user_id | score | method 1 | 10.0 | drop 3 | 30.0 | drop 5 | 50.0 | drop 6 | 60.0 | drop 1 | 10.0 | ffill 2 | 10.0 | ffill 3 | 30.0 | ffill 4 | 30.0 | ffill 5 | 50.0 | ffill 6 | 60.0 | ffill 1 | 10.0 | mean 2 | 37.5 | mean 3 | 30.0 | mean 4 | 37.5 | mean 5 | 50.0 | mean 6 | 60.0 | mean Mean of non-null scores is 37.5; ffill carries prior values.
Topics: lakebench, pandas, fillna, dropna, ffill.
More interview problems · All interview problems · Learn data engineering
Interview-style drill: Show dropna, mean fill, and ffill strategies labeled by method.
From `df`, build three frames: (1) dropna rows, (2) fill score with column mean, (3) forward-fill score. Concat with a `method` column in {drop,mean,ffill}. Sort by `method`, `user_id`. Assign `result`.