Pandas data engineering interview problem. Difficulty: beginner. Pattern: Merges. About 12 minutes. Free to practice.
Demonstrate inner/left/right/outer merges stacked with a join_type column. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Merge df1 and df2 on user_id with how in {inner,left,right,outer}. Concat the four results with a join_type column. Sort by join_type, user_id. Assign result.
Input: df1/df2 Output: user_id | spend | segment | join_type 2 | 20.0 | A | inner 3 | 30.0 | B | inner 1 | 10.0 | None | left 2 | 20.0 | A | left 3 | 30.0 | B | left 2 | 20.0 | A | outer 3 | 30.0 | B | outer 1 | 10.0 | None | outer 4 | None | C | outer 2 | 20.0 | A | right 3 | 30.0 | B | right 4 | None | C | right Each how produces a different null pattern; exact row order follows sort.
Topics: lakebench, pandas, merge, joins.
More interview problems · All interview problems · Learn data engineering
Interview-style drill: Demonstrate inner/left/right/outer merges stacked with a join_type column.
Merge `df1` and `df2` on `user_id` with how in {inner,left,right,outer}. Concat the four results with a `join_type` column. Sort by `join_type`, `user_id`. Assign `result`.