Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. System Design

Interview prep

System design as a prompt, not a slideshow.

Mid-level data engineering interview prompts. Each one has a Scenario, a teaching Solution, and community Q&A. Free problems include the full solution; Pro unlocks the rest.

  • HardFree75 min

    Real-Time Analytics Pipeline (2M Events/Second)

    Design a pipeline that ingests 2M mobile events/sec (~1KB each), processes them in real time, and serves dashboards with a 10-second refresh SLA.

  • MediumPro55 min

    Batch ETL Pipeline for E-Commerce (2M Orders/Day)

    Design a batch pipeline that processes 2M orders/day and lands warehouse tables by 6 AM for daily reports.

  • HardPro55 min

    Log Aggregation System (10,000 Servers)

    Collect logs from 10k servers at ~250 GB/sec aggregate, support sub-second error search for 7 days, and archive for 1 year.

  • MediumPro55 min

    CDC Pipeline with <5-Minute Latency

    Sync a 500GB PostgreSQL OLTP database to a warehouse with under 5-minute latency, including deletes, schema changes, and reconciliation.

  • HardPro55 min

    Feature Store for ML (Batch + Real-Time)

    Design a feature store for batch training (millions of vectors) and real-time inference (<10ms), with point-in-time correctness.

  • EasyFree75 min

    Lambda vs Kappa Architecture

    Explain Lambda (batch + speed layers) vs Kappa (streaming-only), when to recommend each, and Lambda's code-divergence failure mode.

  • MediumFree75 min

    Medallion Architecture (Bronze, Silver, Gold)

    Design a lakehouse with Bronze (raw), Silver (cleaned), and Gold (business aggregates), including tier boundaries and consumers.

  • MediumPro55 min

    Event Sourcing vs State-Based Storage

    Compare event sourcing (store every state change) vs storing current state only; when the complexity is worth it.

  • HardPro55 min

    Data Mesh Architecture

    Design a federated data platform where domain teams own data products; define central platform duties, interoperability, and governance.

  • MediumPro55 min

    Reverse ETL Architecture

    Push transformed warehouse data into operational tools (e.g. Snowflake → Salesforce) with sync frequency, conflict resolution, and rate limits.

  • MediumPro55 min

    Materialized Views vs Pre-Computed Aggregates

    Compare DB-managed materialized views vs pipeline-managed aggregate tables; refresh storms and when to use each.

  • MediumPro55 min

    Lakehouse vs Traditional Warehouse

    Compare Iceberg/Delta/Hudi lakehouses vs Snowflake/BigQuery/Redshift on latency, ACID, cost, openness, and ML.

  • HardFree75 min

    Real-Time Fraud Detection (<100ms Latency)

    Design fraud detection for ~5k TPS payments with <100ms decision latency, including features, scoring, and feedback loops.

  • HardPro55 min

    IoT Sensor Data Pipeline (1M Devices)

    Streaming analytics for 1M devices (~200KB/s aggregate): rolling averages, anomalies, time-series storage.

  • HardPro55 min

    Sessionization System for User Activity

    Group user events into sessions with a 30-minute inactivity gap; handle out-of-order events, late data, and large keyed state.

  • MediumPro55 min

    Real-Time Dashboard with 10-Second Refresh

    Business metrics with 10s refresh using sliding windows, a fast serving store, and WebSocket push.

  • MediumFree75 min

    Exactly-Once Processing in Streaming

    Explain effectively-exactly-once processing: idempotent writes, at-least-once delivery, transactional APIs, and offset management.

  • MediumPro55 min

    Handle Late-Arriving Data in Streaming

    Strategy for late data: watermarks, allowed lateness, nightly reprocessing, append-only design with read-time dedupe.

  • HardPro55 min

    Batch ETL for 10TB Daily Clickstream

    Process 10TB/day clickstream with Protobuf, Flink sessionization + Spark daily aggregations, and spot-instance cost control.

  • MediumPro55 min

    Data Warehouse for E-Commerce (Star Schema)

    Star schema for ~2M orders/day with facts (orders, clicks, inventory) and SCD Type 2 dimensions.

  • EasyPro55 min

    Slowly Changing Dimensions (SCD Type 1 vs Type 2)

    Explain SCD Type 1 vs Type 2 and how to implement Type 2 in a batch pipeline.

  • MediumPro55 min

    Partitioning Strategy for Large Tables

    Partition a 100B-row fact table: date partitions, clustering, Z-ordering, and avoiding small files.

  • MediumPro55 min

    Backfill Strategy for Historical Data

    Reprocess 2 years after a bug fix using atomic swaps, incremental date ranges, without disrupting live consumers.

  • HardPro55 min

    Cost Optimization for 50TB/Day Pipeline

    Cost-optimize 50TB/day: Parquet/ZSTD, pruning, spot batch, autoscaling stream, hot/warm/cold lifecycle.

  • HardPro55 min

    Data Quality Framework for 500 Tables

    DQ framework for 500 tables / 200 pipelines: YAML contracts, GX/dbt tests, Grafana, circuit breakers.

  • MediumPro55 min

    Schema Evolution Without Downtime

    Evolve schemas safely: registry compatibility, explicit columns (no SELECT *), Delta mergeSchema.

  • MediumPro55 min

    Data Lineage Tracking

    Column-level lineage from source to dashboard for impact analysis and compliance; DataHub/OpenMetadata.

  • MediumPro55 min

    Handle PII Data Securely

    PII controls: encryption, column masking, RBAC, audit trails for GDPR/CCPA.

  • MediumPro55 min

    Data Freshness Monitoring

    Monitor freshness SLAs, volume anomalies, and P1/P2 alerting tiers.

  • MediumPro55 min

    Data Contracts Between Teams

    Contracts between producers and consumers: schema, freshness, volume, quality tests, enforcement, versioning.

  • MediumPro55 min

    Handle Data Skew in Spark Join

    One user_id has 90% of transactions - fix Spark join skew with salting, broadcast joins, and AQE.

  • MediumPro55 min

    Backpressure in Streaming Pipeline

    Explain streaming backpressure and strategies: scale bottleneck, Kafka buffering, signals, load shedding.

  • EasyPro55 min

    Small Files Problem in Lakehouse

    Fix 2000x1MB Parquet files via compaction and prevent with batched writes / Delta auto-optimize.

  • MediumPro55 min

    Monitor Production Pipeline Health

    Five-layer monitoring: job health, freshness, volume, quality, infrastructure (Kafka lag, Spark memory).

  • HardPro55 min

    Disaster Recovery for Multi-Region Pipeline

    DR for multi-region pipelines: cross-region Kafka/S3 replication, failover, RTO/RPO.

  • MediumPro55 min

    Optimize Query Performance in Warehouse

    Speed up a 30-minute query on 10TB: pruning, clustering, MVs, rewrite, warehouse sizing.

  • HardPro55 min

    Hybrid Batch + Streaming Pipeline

    Streaming for sub-minute metrics and batch for backfills/heavy aggs, sharing transformation logic (Kappa-inspired).

  • HardPro55 min

    Event-Driven Order Processing System

    Food delivery order flow with Kafka state transitions, transactional outbox, and timeout cancellations.

  • HardPro55 min

    Multi-Tenant Analytics Platform

    Isolated tenant data with shared compute: schemas/warehouses, per-tenant cost budgets, tenant RBAC.

  • HardPro55 min

    Unified Analytics Platform (BI + ML + Real-Time)

    Unified platform: Snowflake BI, Iceberg ML, ClickHouse real-time; shared transforms, catalog lineage, quality at zone boundaries.

LakeBenchPractice today. Build tomorrow.

Warehouse practice that runs in the tab, not on a cluster. Learn concepts, solve interview drills, and mock the round in one place.

Product

  • Studio sandbox
  • Capstone projects

Practice

  • SQL interview questions
  • PySpark interview questions
  • Python interview questions
  • DE theory questions
  • LeetCode for data engineers

Company

  • About
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds & cancellation
  • Shipping & delivery

© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.

No cluster. No install. Just the tab.