Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Conditional column without apply

Pandas data engineering interview problem. Difficulty: beginner. Pattern: Filtering. About 10 minutes. Part of the Pro drill bank.

Create a price tier column from rules in a vectorized way, handling missing prices. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

df has id and price; price can be missing. Add a tier column: 'budget' when price is below 10 'standard' when price is at least 10 and below 50 'premium' when price is 50 or more 'unknown' when price is missing Columns: id, price, tier. Keep the row order, reset the index. Assign the DataFrame to result. Prefer a vectorized expression over a row-by-row function.

Requirements

  • Four possible labels.

Constraints

  • Boundaries are inclusive on the lower side.
  • Prices are non-negative.

Examples

Input: df id | price 1 | 4.99 2 | 10 3 | 49.99 4 | 50 5 | NULL 6 | 0 Output: id | price | tier 1 | 4.99 | budget 2 | 10 | standard 3 | 49.99 | standard 4 | 50 | premium 5 | NULL | unknown 6 | 0 | budget 10.0 is standard (the lower boundary is inclusive) and 50.0 is premium. The missing price is unknown.

Topics: lakebench, pandas, np.select, vectorized, apply.

More interview problems · All interview problems · Learn data engineering

beginner

Conditional column without apply

Interview-style drill: Create a price tier column from rules in a vectorized way, handling missing prices.

`df` has `id` and `price`; `price` can be missing. Add a `tier` column: - `'budget'` when `price` is below 10 - `'standard'` when `price` is at least 10 and below 50 - `'premium'` when `price` is 50 or more - `'unknown'` when `price` is missing Columns: `id`, `price`, `tier`. Keep the row order, reset the index. Assign the DataFrame to `result`. Prefer a vectorized expression over a row-by-row function.