Interview prep
Mid-level data engineering interview prompts. Each one has a Scenario, a teaching Solution, and community Q&A. Free problems include the full solution; Pro unlocks the rest.
Design a pipeline that ingests 2M mobile events/sec (~1KB each), processes them in real time, and serves dashboards with a 10-second refresh SLA.
Design a batch pipeline that processes 2M orders/day and lands warehouse tables by 6 AM for daily reports.
Collect logs from 10k servers at ~250 GB/sec aggregate, support sub-second error search for 7 days, and archive for 1 year.
Sync a 500GB PostgreSQL OLTP database to a warehouse with under 5-minute latency, including deletes, schema changes, and reconciliation.
Design a feature store for batch training (millions of vectors) and real-time inference (<10ms), with point-in-time correctness.
Explain Lambda (batch + speed layers) vs Kappa (streaming-only), when to recommend each, and Lambda's code-divergence failure mode.
Design a lakehouse with Bronze (raw), Silver (cleaned), and Gold (business aggregates), including tier boundaries and consumers.
Compare event sourcing (store every state change) vs storing current state only; when the complexity is worth it.
Design a federated data platform where domain teams own data products; define central platform duties, interoperability, and governance.
Push transformed warehouse data into operational tools (e.g. Snowflake → Salesforce) with sync frequency, conflict resolution, and rate limits.
Compare DB-managed materialized views vs pipeline-managed aggregate tables; refresh storms and when to use each.
Compare Iceberg/Delta/Hudi lakehouses vs Snowflake/BigQuery/Redshift on latency, ACID, cost, openness, and ML.
Design fraud detection for ~5k TPS payments with <100ms decision latency, including features, scoring, and feedback loops.
Streaming analytics for 1M devices (~200KB/s aggregate): rolling averages, anomalies, time-series storage.
Group user events into sessions with a 30-minute inactivity gap; handle out-of-order events, late data, and large keyed state.
Business metrics with 10s refresh using sliding windows, a fast serving store, and WebSocket push.
Explain effectively-exactly-once processing: idempotent writes, at-least-once delivery, transactional APIs, and offset management.
Strategy for late data: watermarks, allowed lateness, nightly reprocessing, append-only design with read-time dedupe.
Process 10TB/day clickstream with Protobuf, Flink sessionization + Spark daily aggregations, and spot-instance cost control.
Star schema for ~2M orders/day with facts (orders, clicks, inventory) and SCD Type 2 dimensions.
Explain SCD Type 1 vs Type 2 and how to implement Type 2 in a batch pipeline.
Partition a 100B-row fact table: date partitions, clustering, Z-ordering, and avoiding small files.
Reprocess 2 years after a bug fix using atomic swaps, incremental date ranges, without disrupting live consumers.
Cost-optimize 50TB/day: Parquet/ZSTD, pruning, spot batch, autoscaling stream, hot/warm/cold lifecycle.
DQ framework for 500 tables / 200 pipelines: YAML contracts, GX/dbt tests, Grafana, circuit breakers.
Evolve schemas safely: registry compatibility, explicit columns (no SELECT *), Delta mergeSchema.
Column-level lineage from source to dashboard for impact analysis and compliance; DataHub/OpenMetadata.
PII controls: encryption, column masking, RBAC, audit trails for GDPR/CCPA.
Monitor freshness SLAs, volume anomalies, and P1/P2 alerting tiers.
Contracts between producers and consumers: schema, freshness, volume, quality tests, enforcement, versioning.
One user_id has 90% of transactions - fix Spark join skew with salting, broadcast joins, and AQE.
Explain streaming backpressure and strategies: scale bottleneck, Kafka buffering, signals, load shedding.
Fix 2000x1MB Parquet files via compaction and prevent with batched writes / Delta auto-optimize.
Five-layer monitoring: job health, freshness, volume, quality, infrastructure (Kafka lag, Spark memory).
DR for multi-region pipelines: cross-region Kafka/S3 replication, failover, RTO/RPO.
Speed up a 30-minute query on 10TB: pruning, clustering, MVs, rewrite, warehouse sizing.
Streaming for sub-minute metrics and batch for backfills/heavy aggs, sharing transformation logic (Kappa-inspired).
Food delivery order flow with Kafka state transitions, transactional outbox, and timeout cancellations.
Isolated tenant data with shared compute: schemas/warehouses, per-tenant cost budgets, tenant RBAC.
Unified platform: Snowflake BI, Iceberg ML, ClickHouse real-time; shared transforms, catalog lineage, quality at zone boundaries.