The interview setup
Design a cost-optimized pipeline for 50TB/day. Discuss columnar formats (Parquet, ZSTD), partition pruning, spot for batch, autoscaling for streaming, and hot/warm/cold lifecycle.
Start by scaring the room with 50TB × 30 ≈ 1.5 PB/month before discipline. Then apply levers systematically.
Background concepts from first principles
Where cloud data bills come from
Stored bytes, scanned bytes (query engines), compute hours (batch/stream), and egress. Architecture maps to those meters.
Columnar + compression
Parquet with ZSTD (or similar) stores only needed columns for many queries and compresses well. Fat JSON as long-term truth is a cost anti-pattern.
Lifecycle tiers
Hot data on faster/expensive tiers for days. Warm compressed object storage for months. Cold/archive for years if compliance requires. Aggregates retained longer than raw.
Spot and autoscale
Batch on spot with restart safety. Streams autoscale on lag rather than "max peak forever."
Pruning
The cheapest byte is the one never scanned. Partitions and clustering enforce that.
The expanded problem
- Estimate monthly storage under lifecycle policies
- Optimize compute for batch and streaming separately
- Ensure query pruning still works
- Keep critical SLAs while cutting waste
Ingest --> hot (days) --> warm compressed --> cold archive
\-> aggregates retained longer (small)What good looks like
- PB math first
- Concrete levers with numbers
- Separate batch vs stream compute strategy
- Compliance exceptions called out
- SLA-critical versus best-effort paths
Clarifying questions
- Retention by dataset?
- Which pipelines are SLA-critical versus best-effort?
- Query patterns that must stay fast?
- Any regulated "keep raw N years" requirements?
Out of scope
- Negotiating vendor discounts as the whole answer
- Deleting compliance data to hit a savings target
Clarifying questions strategy
Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.
What good looks like
Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.
Scope control
Protect the critical path. Park optional marts after the SLA landing succeeds.
Clarifying questions strategy
Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.
What good looks like
Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.
Scope control
Protect the critical path. Park optional marts after the SLA landing succeeds.
Clarifying questions strategy
Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.
What good looks like
Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.
Scope control
Protect the critical path. Park optional marts after the SLA landing succeeds.
Clarifying questions strategy
Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.
What good looks like
Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.
Scope control
Protect the critical path. Park optional marts after the SLA landing succeeds.
Clarifying questions strategy
Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.
What good looks like
Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.
Scope control
Protect the critical path. Park optional marts after the SLA landing succeeds.
Interview framing details
You are expected to teach while you design. Start from a concrete failure, define the terms, then put the algorithm or layout on the board. Keep the scope tight: one grain, one SLA, one publication method.
Ask only the clarifying questions that would change your diagram. Write assumptions when the interviewer asks you to decide. Call out out-of-scope work so you do not burn the clock on tooling trivia.
A strong close restates the critical path, the failure mode you fear most, and what you would ship in week one.
Interview framing details
You are expected to teach while you design. Start from a concrete failure, define the terms, then put the algorithm or layout on the board. Keep the scope tight: one grain, one SLA, one publication method.
Ask only the clarifying questions that would change your diagram. Write assumptions when the interviewer asks you to decide. Call out out-of-scope work so you do not burn the clock on tooling trivia.
A strong close restates the critical path, the failure mode you fear most, and what you would ship in week one.