Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. The 45-day data engineer plan

45-day sprint

The 45-day data engineer plan

Not “learn data engineering.” The goal is narrower: can you sit in a DE1 or DE2 interview, take a real pipeline problem, design it, code it, explain the trade-offs, debug it, and defend your decisions. Two hours taught a day, plus deliberate practice, across 11 phases and 45 days. One coherent stack instead of twenty tools at twenty percent each.

View full curriculum roadmap →

The Day 45 Benchmark

The Complete Whiteboard Pipeline Walkthrough

17 skills · 45 days · 3 capstones

By Day 45 you should be able to walk through this lifecycle out loud without hesitation: defending table grain, partition keys, compute trade-offs, and failure recovery across every link in the chain. Click any skill to jump directly to its lesson:

d.01SQL
d.08Python
d.13ETL / ELT
d.14Data Modeling
d.14Warehouse
d.16Data Lake
d.18Cloud
d.23PySpark
d.29Orchestration
d.33dbt
d.15CDC / SCD
d.34Data Quality
d.26Performance
d.39Streaming
d.42CI / CD
d.41System Design
d.43Production Troubleshooting
Converging into three production capstone projects instead of sitting as isolated, disconnected tools.
View Capstones

Not sure where to begin?

Answer two questions and we'll point you at one first lesson and one free practice ticket.

Where should I start?Start Day 1 (SQL Foundations)
Personalized Plan RhythmOn pace

45-Day Intensive Track

1 curriculum day daily (4 to 8 hours deliberate practice)

Day-0 Diagnostic
Curriculum Progress
0 / 45(0%)
Next Up for You
Day 1: SQL Foundations
Open Day 1 Workspace →
Projected Finish
2026-11-19
Based on current 45-day pacing
Interview Countdown
Not configured
Target company interview readiness

How each day works

2 hours taught (concept, live coding, interview questions, build something), then 6-8 hours of deliberate practice. Not “go watch 5 hours of YouTube.” Every day ends with a specific engineering assignment.

The daily rhythm: block, time, and what it covers
BlockTimeWhat you do
Concept30mThe idea, in plain language, before any syntax.
Live coding / architecture30mWatch or work through the pattern being built, not just described.
Interview questions30mThe questions this topic actually gets asked in interviews.
Build something30mA small, real example of the day's concept, not a toy snippet.
Rebuild without looking1hClose everything. Rebuild the day's core example from memory, then compare.
Project implementation2hApply today's concept to whichever capstone project you picked.
SQL/Python practice1hDrills unrelated to today's topic, to keep older material warm.
Interview questions1hAnswer today's prompts out loud, not just in your head.
Revision30mRedo one exercise from 1, 3, 7, and 14 days ago from memory.
Engineering journal30mWhat you learned, what you implemented, what broke, why it broke, how you fixed it, what trade-off you made, and which interview questions you couldn't answer.

That is roughly 8 hours a day. The final 30 minutes, the journal, is the one that turns Day 45 into a personal knowledge base.

The 45 days

Day 7, 14, 21, 28, 35, 42 are review days - see the revision system below.

Eleven phases, one coherent stack instead of twenty tools at twenty percent. Each phase links to the closest matching Lakebench track, but the day list below is the exact plan, not a re-sorted version of it.

01SQL Masterclass

Days 1 to 7

SQL is the single skill worth over-investing in: not basic SELECT queries, but difficult interview problems solved without panic.

SQL & Warehousing
  1. Day 1
    SQL FoundationsOpen Day 1 Page & Workspace

    Teaches: SELECT, WHERE, ORDER BY, DISTINCT, LIMIT, aliases, NULL, CASE, CAST, basic functions, why the database executes clauses in a fixed logical order

    Build: customers, orders, products, and payments tables. Solve 30 queries against them.

  2. Day 2
    AggregationOpen Day 2 Page & Workspace

    Teaches: GROUP BY, HAVING, COUNT, COUNT DISTINCT, SUM, AVG, MIN/MAX

    • •Find customers whose spending is above the average customer spend.
    • •Find top 3 products by revenue per category.
    • •Find daily revenue.
  3. Day 3
    JOINSOpen Day 3 Page & Workspace

    Teaches: INNER, LEFT, RIGHT, FULL, CROSS, SELF JOIN

    • •Why did your row count increase after joining two tables? This should become second nature.
  4. Day 4
    Window FunctionsOpen Day 4 Page & Workspace

    Teaches: ROW_NUMBER, RANK, DENSE_RANK, LAG, LEAD, SUM OVER, AVG OVER, PARTITION BY, ORDER BY

    • •Find the second highest salary per department.
    • •Find the previous day's revenue.
    • •Find each customer's first purchase.
    • •Find consecutive transactions.
  5. Day 5
    Advanced SQLOpen Day 5 Page & Workspace

    Teaches: CTEs, recursive CTE concept, subqueries, correlated subqueries, EXISTS, UNION, UNION ALL, INTERSECT, EXCEPT

    • •Find customers whose spending is greater than the average customer spending using CTEs.
    • •Traverse an organizational hierarchy tree with a recursive CTE.
    • •Identify missing records between source and warehouse using EXCEPT and INTERSECT.
    • •Find employees earning more than their department's average with a correlated subquery.
  6. Day 6
    SQL PerformanceOpen Day 6 Page & Workspace

    Teaches: indexes, query plans, EXPLAIN, EXPLAIN ANALYZE, partitioning, clustering concepts, predicate pushdown, avoiding SELECT *, joins and cardinality, materialized views

    Build: Deliberately write a slow query, then optimize it.

    • •Profile query execution plans using EXPLAIN and EXPLAIN ANALYZE.
    • •Compare unindexed table scans against index lookups on 15,000 records.
    • •Apply predicate pushdown and SARGable range filters to optimize slow queries.
    • •Pre-aggregate fact tables before joining to avoid join cardinality row explosions.
  7. Day 7
    SQL Interview DayOpen Day 7 Page & Workspace
    • •Complete the 35-question mock exam: 10 Easy, 10 Medium, 10 Hard, and 5 DE Debugging problems.
    • •Perform the 2-hour mistake review classifying errors by Syntax, Concept, Tool, Grain, and Edge Cases.
    • •Verify solutions interactively against live DuckDB mock datasets with Schema Inspector.
    • •Run the 'Blank Page Test' across Foundations, Joins, Aggregations, Windows, CTEs, and Performance.

    No new concepts. A 3-hour mock SQL exam: 10 easy, 10 medium, 10 hard, and 5 real-world DE problems. Then 2 hours reviewing every mistake. By the end of today, SQL should be interview-ready.

02Python for Data Engineering

Days 8 to 12

Not turning you into a software engineer. The question is narrower: can you write production-quality Python that processes data and builds pipelines?

Core PythonPandas
  1. Day 8
    Python FundamentalsOpen Day 8 Page & Workspace

    Teaches: data types, lists & dicts, hash sets, loops, functions, comprehensions, csv.DictReader

    Build: Dirty CSV ingestion → parsing & validation → clean CSV export.

    • •Complete 7 drills on strings, numeric casting, set deduplication, and accumulators.
    • •Build dirty CSV pipeline handling unparseable numbers and missing columns.
    • •CPython memory model, GIL limits, and O(1) vs O(N) lookup complexity.
    • •Blank page test: implement core ingestion without autocomplete.

    Data engineering Python: safe string parsing, currency precision, O(1) set lookups, and pure functions.

  2. Day 9
    Real PythonOpen Day 9 Page & Workspace

    Teaches: modules, exception handling, logging levels, file descriptors, JSON & CSV, env vars

    Build: API → Python → JSON validation → CSV export with structured logging.

    • •Complete 7 drills on modules, try/except blocks, logging, and DictWriter.
    • •Build order ingestion pipeline filtering invalid and cancelled records.
    • •Simulate 5 failure modes: missing keys, bad floats, broken paths, and missing env vars.
    • •Blank page test: explain the architectural role of each pipeline stage.

    Production habits: targeted exceptions, log levels, file handling, and 12-factor configuration.

  3. Day 10
    PandasOpen Day 10 Page & Workspace

    Teaches: read_csv, read_json, merge, groupby, vectorization, missing values, datetime, aggregation

    Build: orders.csv + customers.json → Pandas clean & merge → aggregate report + data quality summary.

    • •Complete 15 drills on boolean masks, fillna, to_datetime, drop_duplicates, and merge validation.
    • •Build e-commerce pipeline joining orders with customer dimensions and data quality checks.
    • •SQL-to-Pandas translation: group by, joins, and window equivalents.
    • •Blank page test: vectorization vs apply performance and merge fan-out risks.

    Know Pandas for local profiling and wrangling, but prioritize vectorization over slow apply loops.

  4. Day 11
    Production PythonOpen Day 11 Page & Workspace

    Teaches: config management, structured logging, custom exceptions, retries & backoff, idempotency, pytest

    Build: Modular pipeline (extract, transform, load, config, utils) with idempotent upserts and automated tests.

    • •Complete 15 drills on 12-factor configs, structured logs, exponential backoff, and idempotent upserts.
    • •Build modular production pipeline with dedicated extract, transform, and load modules.
    • •Simulate 5 failure injections: HTTP 500s, transient 503s, schema drift, and replay audits.
    • •Blank page test: state recovery and idempotent write guarantees.

    Production pipelines survive network blips, malformed payloads, and repeated Airflow retries safely.

  5. Day 12
    Python Interview DayOpen Day 12 Page & Workspace

    Teaches: REST extraction, schema validation, dead-letter quarantine, pure transforms, idempotent upsert, pytest

    Build: REST API → validation & DLQ quarantine → transform → database upsert with failure injection.

    • •Complete 15 drills on timeouts, two-tier validation, DLQ routing, and upserts.
    • •Build complete day-12 pipeline package with full pytest assertions.
    • •Neutralize 6 failure scenarios: missing keys, invalid types, duplicate streams, and network dropouts.
    • •Defend 10 senior interview questions and complete the Phase 2 blank-page test.

    Phase 2 Capstone: validate schemas, isolate bad rows in quarantine, handle 503 retries, and guarantee idempotency.

03Data Engineering Fundamentals

Days 13 to 17

Now you learn what data engineering actually is, not just the tools around it.

SQL & Warehousing
  1. Day 13
    ETL vs ELT, Batch vs Streaming, and Modern Data PipelinesOpen Day 13 · Pro

    Teaches: ETL vs ELT, batch vs streaming, micro-batching, medallion stages, raw immutable storage, replay recovery

    Build: Multi-source ingestion → immutable partitioned raw storage → dual ETL vs ELT comparison → analytics datamart.

    • •Complete 15 drills on ETL vs ELT decisions, Hive partitioning, tumbling windows, and replayability.
    • •Inspect and run 8 pipeline modules comparing Python ETL to SQL ELT execution.
    • •Simulate 5 production failures: warehouse corruption, silent row drops, schema drift, and rate limits.
    • •Defend 10 senior architecture interview questions and complete the blank-page review.

    System trade-offs: compute costs of pre-load transforms vs in-warehouse SQL, and raw storage as replay insurance.

  2. Day 14
    Data Warehousing: Star Schema, Snowflake, Keys, and GrainOpen Day 14 · Pro

    Teaches: OLTP vs OLAP, fact tables, dimension tables, star vs snowflake, surrogate vs natural keys, declaring grain

    Build: Kimball Star Schema in SQLite (orders, customers, products, dates) with grain fan-out debugger.

    • •Complete 15 drills on OLTP vs OLAP profiling, additive measures, grain statements, and surrogate keys.
    • •Execute Star Schema and Snowflake modules in Pyodide with multi-hop join profiling.
    • •Simulate 5 failure modes: join fan-out inflation, missing -1 dimension members, and grain pollution.
    • •Defend 10 senior interview questions and complete the data modeling blank-page exam.

    Declare the grain in plain English first: grain dictates primary keys, valid metrics, and prevents join fan-out.

  3. Day 15
    Slowly Changing Dimensions (SCD Type 0, 1, & 2)Open Day 15 · Pro

    Teaches: SCD Type 0 (Freeze), SCD Type 1 (Replace), SCD Type 2 (Version), valid_from & valid_to, is_current flag, point-in-time joins

    Build: Multi-batch SCD Type 2 dimension engine with surrogate keys, half-open intervals, and time-travel SQL.

    • •Complete 15 drills on SCD strategy selection, half-open intervals, time-travel queries, and row hashing.
    • •Execute SCD2 customer pipeline with automated temporal invariant testing.
    • •Simulate 5 failure modes: duplicate active versions, inclusive boundary overlap, and late-arriving updates.
    • •Defend 10 senior interview questions and complete the SCD recall test.

    Mental model: Freeze (0), Replace (1), Version (2). Use half-open intervals [start, end) for point-in-time fact joins.

  4. Day 16
    Data LakesOpen Day 16 · Pro

    Teaches: lake vs warehouse vs lakehouse, Medallion architecture, Parquet vs CSV, Hive partitioning, partition pruning, small files compaction

    Build: Medallion Data Lake engine: raw CSV → bronze standardization → silver partitioned Parquet → gold daily sales aggregate.

    • •Complete 15 drills on Medallion contracts, column projection pushdown, Hive paths, and file compaction.
    • •Run Parquet vs CSV benchmarks, partition pruning simulations, and small file diagnostic modules.
    • •Simulate 5 failure modes: unpartitioned data swamp, high-cardinality partition explosion, and bill spikes.
    • •Defend 10 senior interview questions and complete the storage blank-page review.

    Never destroy raw data. Store columnar Parquet with coarse temporal partitions so engines skip 90%+ of files.

  5. Day 17
    Pipeline Architecture: The Capstone Architecture ChallengeOpen Day 17 · Pro

    Teaches: multi-source integration, read-replica CDC, API exponential backoff & DLQ, Kafka stream dedup, dual-speed SLAs, pipeline idempotency

    Build: Enterprise multi-source platform: PostgreSQL + flaky REST API (with DLQ) + Kafka clickstream → S3 Parquet lake + SCD2 warehouse.

    • •Complete 15 drills on CDC vs watermarks, jittered backoff, stream deduplication, and atomic staging swaps.
    • •Execute end-to-end multi-source pipeline architecture in Pyodide WASM.
    • •Simulate 5 production failures: connection exhaustion, API 429 stampedes, small files, and streaming race conditions.
    • •Defend 10 senior architecture interview scenarios and pass the blank-page challenge.

    Phase 3 Capstone: design first, hit architectural failure modes, correct the design, and defend trade-offs.

04Cloud

Days 18 to 22

Cloud, now that the fundamentals are in place. Same 5 days, same shape - pick AWS, GCP, or Azure below.

Cloud Platforms
  1. Day 18
    AWS FundamentalsOpen Day 18 · Pro

    Teaches: IAM roles & policies, STS temporary tokens, least privilege, S3 object model, EC2 vs Lambda, CloudWatch alarms

    Build: AWS security simulator: IAM evaluation, temporary STS token assumption, and CloudWatch metric alerting.

    • •Complete 15 drills on IAM policies, STS assumption, S3 permissions, and Lambda timeouts.
    • •Run 6 Python modules simulating IAM evaluation, EC2 vs Lambda selection, and CloudWatch.
    • •Simulate 5 production failures: leaked AKIA keys, S3 403 ListBucket denial, and Lambda timeouts.
    • •Defend 10 senior AWS security questions and complete the blank-page exam.

    Never hardcode AWS keys. Attach IAM roles to compute and issue short-lived temporary tokens via STS.

  2. Day 19
    S3 + Data LakeOpen Day 19 · Pro

    Teaches: S3 prefixes vs directories, Medallion zones (Raw/Silver/Gold), Hive partitioning, partition pruning, lifecycle tiering, compaction

    Build: S3 Medallion Lake engine: raw ingestion → schema validation → Hive partitioned Parquet → curated sales mart.

    • •Complete 15 drills on Medallion contracts, Hive paths, Athena scan pruning, and lifecycle decay.
    • •Execute 6 Python modules simulating lake ingestion, partition pruning, and Delete Marker recovery.
    • •Simulate 5 production failures: small files meltdown, duplicate retries, and unpartitioned scans.
    • •Defend 10 senior storage interview questions and pass the blank-page challenge.

    Partition by query access patterns, keep immutable raw data for replay, and automate lifecycle tiering to Glacier.

  3. Day 20
    Athena + GlueOpen Day 20 · Pro

    Teaches: S3 + Glue + Athena flow, Glue Data Catalog, external tables & SerDes, Athena scan pricing ($5/TB), partition pruning, schema evolution

    Build: Serverless query engine: Glue catalog discovery → partition pruning benchmarks → pre-warehouse data quality validation.

    • •Complete 15 drills on Glue catalog schemas, Hive SerDes, scan pricing formulas, and schema drift.
    • •Execute 6 Python modules simulating Athena query orchestration and schema evolution.
    • •Simulate 5 production failures: runaway $18k scan invoice, silent type drift, and full scan timeouts.
    • •Defend 10 senior serverless query questions and pass the blank-page challenge.

    Athena is serverless compute charging by data scanned. Pair Hive-partitioned Parquet with column pruning to cut 99% of costs.

  4. Day 21
    RedshiftOpen Day 21 · Pro

    Teaches: Redshift MPP architecture, DISTSTYLE (KEY, ALL, EVEN), Sort Keys & Zone Maps, columnar compression, COPY from S3, atomic staging merges

    Build: Redshift MPP warehouse engine: parallel S3 COPY → staging deduplication → atomic idempotent merge transaction.

    • •Complete 15 drills on MPP slices, distribution styles, sort key pruning, and COPY options.
    • •Execute 6 Python modules simulating Redshift ingestion, distribution skew, and idempotent merges.
    • •Simulate 5 production failures: single-row INSERT lockups, duplicate PKs on retry, and slice data skew.
    • •Defend 10 senior data warehouse questions and pass the blank-page challenge.

    Redshift doesn't enforce primary key constraints on write. Guarantee idempotency using staging tables and atomic transactions.

  5. Day 22
    AWS PipelineOpen Day 22 · Pro

    Teaches: complete AWS platform flow, high-watermark extraction, S3 lakehouse staging, Glue catalog registration, Redshift COPY & merge, reconciliation

    Build: End-to-end AWS data platform: PostgreSQL incremental extraction → S3 raw/curated → Athena validation → Redshift merge.

    • •Complete 15 drills on watermark tracking, S3 layout, external tables, and warehouse reconciliation.
    • •Execute 6 Python modules simulating the complete AWS pipeline and discrepancy debugger.
    • •Simulate 5 production failures: premature watermark advance, duplicate retries, and missing dimensions.
    • •Defend 10 senior cloud pipeline interview questions and pass the blank-page exam.

    Decouple Postgres from Redshift using S3. Advance watermarks only after warehouse commits, and use LEFT JOIN fallbacks.

05Spark / PySpark

Days 23 to 28

This is where you become much more employable. Current DE-II postings repeatedly emphasize Spark and scalable pipeline processing.

PySpark
  1. Day 23
    Spark FundamentalsOpen Day 23 · Pro

    Teaches: horizontal scaling, driver vs executor, Catalyst DAG, transformations vs actions, narrow vs wide stages, fault tolerance lineage

    Build: Distributed execution simulator: driver scheduling, narrow transformations, shuffle partitions, and worker failure recovery.

    • •Complete 15 drills on driver vs executor roles, narrow vs wide transforms, and shuffle boundaries.
    • •Run 6 Python modules simulating execution stages, partition byte-splits, and lineage recovery.
    • •Simulate 5 production outages: driver OOM via collect(), 10k small files, and shuffle fetch failures.
    • •Defend 10 senior Spark architecture questions and pass the blank-page challenge.

    Spark coordinates distributed compute. Never run df.collect() on large data, and remember files in S3 are not partitions.

  2. Day 24
    PySpark DataFramesOpen Day 24 · Pro

    Teaches: lazy evaluation, the 9 core DataFrame ops, predicate pushdown, bitwise filtering (&, |), avoiding withColumn loops, window deduplication

    Build: E-commerce PySpark pipeline: raw CSV ingestion → data quality filtering → multi-metric aggregation → master data join.

    • •Complete 15 drills on column projections, multi-condition filters, aggregations, and joins.
    • •Run 6 Python modules comparing native expressions to UDFs and testing DataFrame invariants.
    • •Simulate 5 production outages: plan depth StackOverflow, ambiguous column collisions, and UDF bottlenecks.
    • •Defend 10 senior PySpark DataFrame questions and pass the blank-page challenge.

    Never use Python 'and'/'or' on Spark Columns: use '&' and '|'. Avoid chaining withColumn in loops or using df.distinct() blindly.

  3. Day 25
    Spark TransformationsOpen Day 25 · Pro

    Teaches: partition boundaries, narrow vs wide transformations, shuffle write & fetch, repartition vs coalesce, coalesce(1) single-core hazard, partition sizing

    Build: Multi-stage pipeline: narrow filtering → wide shuffle aggregation → hash repartitioning by customer key → coalesce(2) lake write.

    • •Complete 15 drills on narrow vs wide operations, physical execution plans, and shuffle costs.
    • •Run 6 Python modules simulating repartition vs coalesce mechanics and key co-location.
    • •Simulate 5 production outages: coalesce(1) bottleneck, 200 tiny files, and Cartesian join disk thrashing.
    • •Defend 10 senior Spark transformation scenarios and pass the blank-page exam.

    Never use coalesce(1) before writing large datasets: it collapses upstream execution to a single core and crashes workers.

  4. Day 26
    Spark PerformanceOpen Day 26 · Pro

    Teaches: BroadcastHashJoin (<100MB), partition pruning, Parquet predicate pushdown, when to cache vs unpersist, data skew key salting, small files coalesce

    Build: Refactor a slow Spark job with 6 bottlenecks (wide scan, missing broadcast, shuffle skew, coalesce(1)) into a 95% faster pipeline.

    • •Complete 15 drills on broadcast hash joins, partition pruning, predicate pushdown, and key salting.
    • •Run 6 Python modules comparing slow vs optimized execution plans and skew remediation.
    • •Simulate 5 production outages: broadcast join OOM, 4-hour NULL join skew, and blind caching GC freezes.
    • •Defend 10 senior performance scenarios and pass the blank-page optimization challenge.

    The fastest data is data Spark never reads or moves. Use broadcast joins for small dimensions and salt skewed keys.

  5. Day 27
    PySpark Production PipelineOpen Day 27 · Pro

    Teaches: data contracts & early validation, string & enum sanitation, windowed latest-record deduplication, join cardinality diagnostics, pre-write circuit breakers, idempotent writes

    Build: End-to-end PySpark pipeline: raw ingestion → contract validation → window deduplication → broadcast join → pre-write assertions.

    • •Complete 15 drills on schema contracts, window deduplication, join cardinality checks, and idempotency.
    • •Run 6 Python modules validating pipeline stages, cardinality audits, and pre-write assertions.
    • •Simulate 5 production outages: revenue explosion from dimension duplicates, partial write appends, and empty outputs.
    • •Defend 10 senior pipeline engineering scenarios and complete the architecture test.

    Verify dimension uniqueness before joining to prevent row explosion, and write with mode('overwrite') to ensure rerun safety.

  6. Day 28
    Spark Interview DayOpen Day 28 · Pro

    Teaches: Spark execution mental model, Spark UI triage framework, shuffle mechanics & skew, broadcast trade-offs, repartition vs coalesce, executor failure triage

    Build: Spark Interview Lab: 5 TB skew autopsy with salting, shuffle diagnostic analyzer, and small files compaction optimizer.

    • •Defend the 7 mandatory Spark interview areas (slow jobs, shuffle, skew, broadcast, repartition, small files, executor death).
    • •Solve the 5 TB production skew incident (1.7B rows on a single customer) and 4.2M small files outage.
    • •Run 6 Python modules simulating skew salting, shuffle diagnostics, and broadcast analyzers.
    • •Pass the 8 capstone interview criteria and master the 8 blank-page verification questions.

    Never just say 'add more executors' when a job is slow. Triage the Spark UI: find the slow stage, inspect task skew, and check shuffle volume.

06Airflow + Orchestration

Days 29 to 32

Orchestration coordinates the moving parts of modern data systems. Build production DAGs, handle backfills safely, prevent worker starvation, and master failure triage.

Orchestration
  1. Day 29
    Airflow Fundamentals & DAGsOpen Day 29 · Pro

    Teaches: DAGs as directed acyclic graphs, Scheduler vs Webserver vs Executor vs Worker, Task state lifecycle & cascades, Operators vs Sensors vs TaskFlow API, Bitshift dependency chaining (>>), XCom mechanics & metadata DB hazards

    Build: Airflow execution engine: DAG registry, topological sort dependency resolution, state transition simulation, and 48KB XCom overflow protection.

    • •Complete 15 drills on DAG dependencies, cycle detection, state cascades, and XCom pointers.
    • •Run 6 Python modules simulating DAG dependency resolution, execution states, and XCom payload boundaries.
    • •Simulate 5 production failures: scheduler top-level parse death, 500MB XCom DB crash, cycle deadlock, and runaway backfill.
    • •Defend 10 senior Airflow architecture questions and complete the blank-page challenge.

    Never pass dataframes or raw datasets through XCom. Use XCom strictly for lightweight metadata and remote pointers (S3/GCS paths).

  2. Day 30
    Production Airflow & ReliabilityOpen Day 30 · Pro

    Teaches: Logical date vs Execution date vs Data interval, catchup=False vs catchup=True, CLI backfills & idempotent rerun safety, Sensor poke vs reschedule mode, Worker slot starvation prevention, Retries with exponential backoff & jitter, Avoiding top-level Variable.get() / Connection.get() DB blasts

    Build: Production Airflow simulator: logical date intervals, sensor slot starvation remediation, exponential retry jitter, and execution_timeout circuit breakers.

    • •Complete 15 drills on logical date calculations, sensor modes, backoff jitter, and metadata DB protection.
    • •Run 6 Python modules simulating worker slot contention, backfill reconciliation, and variable caching.
    • •Simulate 5 production failures: Celery worker starvation, duplicate date execution, thundering herd retries, and scheduler DB connection pool collapse.
    • •Defend 10 senior production Airflow questions and pass the blank-page challenge.

    Never use sensor poke mode for long waits (>5 min): it locks worker slots and starves all other pipelines. Always use mode='reschedule'.

  3. Day 31
    Airflow + Spark + Cloud LakehouseOpen Day 31 · Pro

    Teaches: Separation of concerns: Airflow orchestrates, Spark computes, Why running PySpark code inside Airflow workers is a critical anti-pattern, SparkSubmitOperator vs EMR / Dataproc / Databricks operators, Deferrable operators & asynchronous triggers, S3/GCS staging handoff: Extract → Raw → Spark → Curated → Warehouse COPY/MERGE, Atomic partition replacement & staging cleanup

    Build: Cloud lakehouse orchestrator: Airflow async trigger DAG → Dataproc/EMR Spark execution → S3 Parquet validation → atomic warehouse MERGE gate.

    • •Complete 15 drills on cloud operator patterns, async deferrable triggers, S3 partition handoffs, and warehouse ingestion gates.
    • •Run 6 Python modules simulating deferrable polling, partition manifest swaps, and staging garbage collection.
    • •Simulate 5 production failures: Airflow worker OOM from local PySpark, cluster timeout stalls, dirty partition loads, and zombie worker slots.
    • •Defend 10 senior cloud orchestration interview questions and complete the blank-page exam.

    Never execute heavy compute on Airflow worker nodes. Treat Airflow strictly as the air traffic controller dispatching jobs to managed clusters.

  4. Day 32
    Orchestration Failure Injection & TriageOpen Day 32 · Pro

    Teaches: The 6 classic production orchestration failures, API rate limits (429) & exponential jitter backoff, Cloud storage eventual consistency & missing partition races, Source schema drift & dead-letter quarantine, Database connection pool exhaustion & connection pooling, Spark executor OOM retry cascade prevention, Metadata DB lock contention triage

    Build: Production Incident Triage Console: live failure injection (429 rate limit, schema drift, DB pool exhaustion, storage lag) with diagnostic logs, real-time remediation, and health verification.

    • •Complete 15 drills on failure classification, rate limit throttling, dead-letter routing, and DB connection limits.
    • •Run 6 Python modules simulating live incident triage, backoff algorithms, and schema drift quarantine.
    • •Solve 6 real-world production incident autopsies with root cause timeline, immediate fix, and architectural prevention.
    • •Defend 10 senior orchestration triage interview scenarios and pass the blank-page triage exam.

    When a pipeline fails, never blindly click 'Clear' to retry. First identify whether the failure is transient, systemic, or a data-corruption hazard.

07dbt + Data Quality

Days 33 to 35

Analytics engineering turns raw data lakes into trusted, tested data models. Master SQL-first transformations, modular DAG refs, data contracts, and watermarked incremental processing.

dbt & AnalyticsData Quality
  1. Day 33
    dbt Core & ModelingOpen Day 33 · Pro

    Teaches: dbt architecture & compilation to native SQL, Sources, staging models & column renaming, ref() DAG generation & lineage, Generic tests (not_null, unique, accepted_values, relationships), dbt macros & Jinja templating, SCD Type 2 snapshots (check vs timestamp strategy)

    Build: dbt transformation engine: raw sources → staging clean views → dimensional fact mart with automated lineage graph compilation and snapshot history tracking.

    • •Complete 15 drills on dbt ref() lineage, source contracts, custom test macros, and snapshot invalidations.
    • •Run 6 Python modules simulating Jinja compilation, DAG topological sorting, and snapshot row mutation.
    • •Simulate 5 production failures: circular ref deadlock, schema drift break, snapshot timestamp drift, and failed uniqueness test.
    • •Defend 10 senior dbt analytics engineering questions and pass the blank-page challenge.

    Never hardcode database or schema names in SQL. Always use {{ source() }} and {{ ref() }} to guarantee atomic DAG compilation and environment isolation.

  2. Day 34
    Data Quality & Anomaly DetectionOpen Day 34 · Pro

    Teaches: Defense in depth: Pre-load vs In-flight vs Post-load verification, Column invariants (uniqueness, referential integrity, range bounds), Row count anomaly detection & statistical deviation (Z-score), Freshness SLAs & silent pipeline stalls, Circuit breakers: quarantine vs pipeline abort, Data contract enforcement

    Build: Data quality & circuit breaker engine: schema enforcement → statistical volume anomaly detector → referential integrity validator → quarantine isolation router.

    • •Complete 15 drills on contract validation, foreign key orphaned checks, statistical anomaly detection, and quarantine routing.
    • •Run 6 Python modules testing circuit breaker trip points, Z-score thresholds, and automated quarantine generation.
    • •Simulate 5 production failures: silent 0-row load, duplicate primary keys, foreign key orphan explosion, and 30% revenue drop anomaly.
    • •Defend 10 senior data quality questions and complete the blank-page exam.

    A failing pipeline that halts execution is an inconvenience; a silent corrupt pipeline that writes dirty metrics to production is an executive catastrophe.

  3. Day 35
    Incremental Processing & WatermarksOpen Day 35 · Pro

    Teaches: Full refresh vs incremental processing trade-offs, Watermark tracking & high-watermark state stores, Lookback windows for late-arriving records, Atomic MERGE & upsert mechanics (is_incremental() macro), Unique key deduplication before warehouse write, Handling hard deletes and tombstones

    Build: Incremental lakehouse pipeline: high-watermark state tracking → late-arriving event lookback buffer → idempotent atomic MERGE → audit log reconciliation.

    • •Complete 15 drills on high-watermark queries, lookback window sizing, MERGE condition optimization, and delete tombstoning.
    • •Run 6 Python modules simulating watermark state transitions, late-record reconciliation, and idempotent rerun guarantees.
    • •Simulate 5 production failures: premature watermark advance, missing lookback window, duplicate rows from retry, and tombstone desync.
    • •Defend 10 senior incremental processing questions and pass the blank-page challenge.

    Always include a lookback window (e.g. current_watermark - interval '3 hours') in incremental filters to catch out-of-order and late-arriving records.

08Project #1 - E-Commerce Data Platform

Days 36 to 38

A complete, production-grade batch data platform: PostgreSQL extraction → Airflow orchestration → S3 Bronze staging → PySpark Silver cleaning & Gold aggregation → Redshift/BigQuery warehouse → dbt metrics → SLA alerting.

  1. Day 36
    Project #1: Ingestion & Lake StagingOpen Day 36 · Pro

    Teaches: Transactional database CDC vs batch extraction, High-watermark state persistence in DynamoDB/PostgreSQL, Partitioned S3/GCS Bronze raw lake staging, Schema validation & payload checksum hashing, REST API event enrichment ingestion, Idempotent extraction chunking

    Build: Platform Stage 1: PostgreSQL & Stripe API batch extractor → checksum validator → partitioned Hive Parquet S3 Bronze lake writer.

    • •Complete 15 drills on chunked extraction, watermark commits, schema hashing, and S3 staging paths.
    • •Run 6 Python modules testing extraction resilience, checksum verification, and network failure retries.
    • •Simulate 5 production failures: database replica timeout, partial chunk crash, schema column drop, and corrupted payload hash.
    • •Defend 10 senior ingestion architecture questions and pass the blank-page challenge.

    Never perform transformations during the raw extraction phase. Write Bronze data in its raw, unmodified form to enable historical replay.

  2. Day 37
    Project #1: PySpark Silver & Gold CurationsOpen Day 37 · Pro

    Teaches: Medallion curation: Bronze raw to Silver cleaned to Gold aggregates, Windowed deduplication by primary key and updated_at timestamp, Data contract casting & null substitution policies, Broadcast dimension enrichment, Gold daily revenue & customer lifetime value fact tables, Atomic S3 partition swaps

    Build: Platform Stage 2: PySpark Silver deduplication & validation engine → Gold metric aggregations → broadcast join enricher → atomic lakehouse writer.

    • •Complete 15 drills on Silver cleaning, window deduplication, broadcast join thresholds, and Gold rollup metrics.
    • •Run 6 Python modules verifying Silver contract cleanliness, Gold revenue reconciliations, and skew handling.
    • •Simulate 5 production failures: dimension duplicate row explosion, skewed seller bottleneck, float precision rounding drift, and partition write crash.
    • •Defend 10 senior Spark curation questions and complete the blank-page exam.

    Always deduplicate Silver data using row_number() over (partition by id order by updated_at desc) before performing any join operations.

  3. Day 38
    Project #1: Airflow, dbt & SLA DefenseOpen Day 38 · Pro

    Teaches: End-to-end Airflow DAG tying extraction, Spark, and dbt together, Warehouse atomic MERGE / COPY ingestion gates, dbt data tests as blocking deployment gates, SLA monitoring & PagerDuty alert routing, Complete system architecture verbal defense, Cost & performance trade-off justification

    Build: Platform Stage 3: End-to-end Airflow master DAG orchestrating extraction, EMR Spark, Redshift COPY, dbt tests, and Slack/PagerDuty webhook alerts.

    • •Complete 15 drills on master DAG dependencies, warehouse copy gates, blocking dbt test triggers, and SLA triage.
    • •Run 6 Python modules simulating end-to-end pipeline execution, reconciliation auditing, and alert dispatch.
    • •Simulate 5 production failures: warehouse lock timeout, dbt test rejection, SLA breach alert, and partial pipeline recovery.
    • •Deliver the 20-minute architecture defense presentation and complete the full blank-page architecture blueprint.

    In a Senior DE interview, walk through this platform using the 'Requirements → Bottlenecks → Trade-offs → Failure Modes' framework.

09Streaming + CDC

Days 39 to 40

Modern data platforms cannot afford 24-hour batch latency. Master real-time change data capture, distributed log queues, event ordering, and streaming state engines.

Streaming
  1. Day 39
    Apache Kafka Architecture & MechanicsOpen Day 39 · Pro

    Teaches: Distributed commit log architecture, Brokers, topics, partitions & segments, Message keys & deterministic partition routing, Producers: acks (0, 1, all) & idempotence, Consumer groups, rebalancing & offset commit semantics, Lag monitoring & consumer scale limits

    Build: Kafka streaming simulator: producer partition router with hashing, multi-consumer group coordinator, offset commit manager, and lag spike monitor.

    • •Complete 15 drills on partition hashing, producer acks trade-offs, consumer group rebalances, and lag calculations.
    • •Run 6 Python modules simulating partition assignment, rebalance stop-the-world pauses, and offset commits.
    • •Simulate 5 production failures: consumer lag explosion, hot partition skew from null keys, duplicate delivery on rebalance, and broker disk exhaustion.
    • •Defend 10 senior Kafka architecture questions and pass the blank-page challenge.

    Partition count is the fundamental unit of parallelism. You cannot have more active consumers in a consumer group than partitions in a topic.

  2. Day 40
    CDC, Debezium & Streaming JoinsOpen Day 40 · Pro

    Teaches: Change Data Capture (CDC) via database write-ahead log (WAL / binlog), Debezium connector architecture & event envelopes (before, after, op), Out-of-order events & watermark event-time processing, Deduplication & exactly-once processing guarantees, Streaming stateful joins (stream-stream vs stream-table), Changelog table compaction

    Build: CDC stream processing engine: PostgreSQL WAL log parser → Debezium event envelope flattener → out-of-order watermark buffer → compacted lakehouse sink.

    • •Complete 15 drills on Debezium envelopes, WAL extraction, watermark windows, and stream deduplication.
    • •Run 6 Python modules simulating out-of-order event arrival, watermark triggers, and log compaction.
    • •Simulate 5 production failures: database WAL disk full from slow consumer, late-arriving event dropped, tombstone loss, and duplicate CDC replays.
    • •Defend 10 senior CDC & streaming questions and complete the blank-page exam.

    Never query source databases directly for real-time changes. Read the database WAL via Debezium to avoid impacting transactional query performance.

10DE2 Engineering

Days 41 to 43

What separates a junior coder from a Senior Staff Data Engineer: resilient distributed systems design, airtight production operations, CI/CD, and live incident triage.

System DesignInfra Practices
  1. Day 41
    DE System Design at ScaleOpen Day 41 · Pro

    Teaches: 500M events/day data platform blueprint, Throughput, storage & compute capacity sizing, Lambda vs Kappa vs Delta lakehouse architectures, Decoupling ingestion, processing, and analytical serving, Multi-tenant isolation & noisy neighbor prevention, Scaling from 500M to 5B events/day: 10x bottlenecks

    Build: System Design interactive sizing calculator: daily event volume → throughput/sec → storage decay → worker cluster sizing → cost estimation model.

    • •Complete 15 drills on throughput calculations, storage IOPS sizing, network egress estimation, and multi-region failover.
    • •Run 6 Python modules simulating peak traffic spikes, partition scaling, and storage cost projections.
    • •Analyze 5 production design trade-offs: Kafka vs SQS, Parquet vs Delta, Redshift vs Snowflake, and Batch vs Streaming.
    • •Defend the complete 500M events/day system design interview defense and pass the blank-page architecture challenge.

    In system design interviews, always start with requirements, SLA numbers, and back-of-the-envelope calculations before drawing single boxes or choosing technologies.

  2. Day 42
    Production Engineering & CI/CDOpen Day 42 · Pro

    Teaches: Git trunk-based development & feature branching, Data pipeline CI/CD with automated testing stages, Docker containerization for reproducible Spark/Airflow environments, Secrets management & credential rotation (IAM, Vault, AWS Secrets Manager), Infrastructure as Code (Terraform) principles, Observability: Metrics, logs, traces, and SLO/SLA error budgets

    Build: Production CI/CD & validation pipeline: GitHub Actions simulator → linting → SQL fluff → PySpark unit tests → Docker build → staging deployment gate.

    • •Complete 15 drills on Git workflows, Dockerfile multi-stage builds, secrets rotation, and error budget calculations.
    • •Run 6 Python modules simulating CI test runners, container health checks, and secret masking.
    • •Simulate 5 production failures: leaked AWS key in git commit, failed deployment rollback, Docker image drift, and breached error budget.
    • •Defend 10 senior production engineering questions and complete the blank-page exam.

    Never test data pipelines in production. Enforce automated pre-merge testing with mock datasets to catch bugs before they corrupt production tables.

  3. Day 43
    Advanced Incident Triage & Root CauseOpen Day 43 · Pro

    Teaches: The 6 mission-critical production incident drills, Scenario 1: Pipeline takes 3 hours instead of 20 mins, Scenario 2: Spark job sudden executor OOM crash, Scenario 3: Financial revenue metric jumps 30% unexpectedly, Scenario 4: Pipeline reports success but destination table is empty, Scenario 5: Duplicate records appear in downstream mart, Scenario 6: Upstream source introduces unannounced schema drift

    Build: Production Incident War Room: real-time incident simulator where student investigates symptoms, queries telemetry, diagnoses root cause, and applies permanent fix.

    • •Complete 15 drills on incident triage order, telemetry analysis, root cause deduction, and blameless post-mortems.
    • •Run 6 Python modules simulating incident diagnostic workflows and verification tests.
    • •Resolve all 6 production emergency scenarios with complete post-mortem writeups.
    • •Defend 10 senior incident management interview questions and pass the blank-page triage challenge.

    Treat every incident as a systemic failure, not human error. Document the root cause timeline and implement automated regression tests to prevent recurrence.

11Capstone

Days 44 to 45

The culmination of your 45-day journey. Build an enterprise-grade real-time streaming platform from scratch, then defend your knowledge across 7 rigorous mock interview rounds.

  1. Day 44
    Capstone Project #2: Real-Time PlatformOpen Day 44 · Pro

    Teaches: Real-time e-commerce architecture specification, PostgreSQL transactional source → Debezium CDC → Kafka event bus, PySpark Structured Streaming with watermarks & stateful deduplication, S3/ADLS Medallion lakehouse (Delta/Iceberg/Parquet), Warehouse dimension modeling & incremental dbt mart, Production monitoring, alerting & end-to-end data contracts

    Build: Capstone Platform: PostgreSQL CDC → Kafka streaming pipeline → PySpark streaming enrichment → Medallion lakehouse → dbt analytical warehouse.

    • •Complete 15 drills on end-to-end contract testing, streaming checkpoints, compaction schedules, and warehouse reconciliation.
    • •Run 6 Python modules verifying streaming state recovery, watermark cutoff adherence, and end-to-end latency.
    • •Simulate 5 production failures: broker disconnect, checkpoint directory corruption, schema drift event, and warehouse connection timeout.
    • •Complete the 100% end-to-end capstone platform checklist and code verification.

    Build every layer independently and verify contracts at each boundary. If a bug occurs, isolate whether the source is CDC, Kafka, Spark, or the warehouse.

  2. Day 45
    7-Round Mock DE Interview DefenseOpen Day 45 · Pro

    Teaches: The 7 rounds of senior data engineering hiring, Round 1: Advanced SQL & Data Modeling (60 min), Round 2: Python Data Structures & Core Algorithms (45 min), Round 3: PySpark & Distributed Computing Deep Dive (45 min), Round 4: Data Engineering Pipeline Architecture & Storage (60 min), Round 5: Distributed System Design at 500M+ Scale (60 min), Round 6: Project Portfolio Deep Dive & Battle Scars (60 min), Round 7: Senior Staff Behavioral & Leadership Defense (30 min)

    Build: Full 7-Round Interview Simulator: interactive question prompt, rubric evaluation, weak answer red flags, model senior staff responses, and scoring scorecard.

    • •Complete all 7 mock interview rounds with live defense questions and scoring rubrics.
    • •Review 21 red flag traps that cause immediate rejections in Senior DE interviews.
    • •Run 6 Python modules demonstrating live coding solutions for SQL, Python, and Spark technical rounds.
    • •Master the 10 capstone graduation questions and complete the final 45-day Senior Data Engineer Certification Exam.

    Senior engineers don't win interviews by reciting syntax. They win by explaining trade-offs, anticipating edge cases, and showing proven production empathy.

The 3 capstone projects

Pick one domain and build all three against it so the story stays coherent end to end. These map onto the same capstones you already write inside each Learn track, described the way an interviewer would ask about them. The full list, with links to every capstone lesson, lives on /projects.

Project A

E-commerce Analytics Platform

The primary batch project. Everything else references it.

Sources: PostgreSQL, REST API, CSV. Features: incremental loads, SCD2, data quality, partitioning, deduplication, schema validation, logging, monitoring.

Tables: customers, products, orders, order_items, payments, shipments, returns, reviews

Core PythonSQL & WarehousingPandasPySparkdbt & Analytics
Project B

Real-Time Customer Behavior Platform

Same e-commerce domain as A, so the two connect. The streaming project.

Application → Kafka → Spark Streaming → S3 → Redshift. Teaches partitions, offsets, consumer groups, event ordering, late events, duplicates, and watermarks.

Tables: page_view, product_view, add_to_cart, checkout, purchase

The queue is taught and practiced as a mocked broker in Python, per the Streaming track. No Kafka process runs in this tab.

StreamingPySpark
Project C

Customer 360 / CDC Platform

The project for the strongest DE2-depth conversations.

Operational DB → CDC → Kafka → Spark → S3 → Warehouse → dbt → Customer 360. Implements SCD2, CDC, MERGE, incremental processing, data quality, and schema evolution.

Tables: customer, customer_address, customer_orders, customer_payments, customer_interactions

OrchestrationSQL & Warehousing

Do not invent a job history for these

Do not claim you built these systems at an employer when you did not. Interviewers drill: how many records, what was your SLA, what incident happened, who were your consumers, what was your AWS bill, how did you handle on-call. A fabricated story eventually collapses.

Instead, build the three projects above to a level where they genuinely resemble production systems, put them in a public GitHub repo with an architecture diagram, README, sample data, tests, and documented trade-offs, and describe them accurately as hands-on capstone work. If you already worked at a company for a year or two, the stronger move is to retrofit what you genuinely learned there onto the real business context you already know, so a 45-minute deep dive draws on things you actually understand rather than a memorized lie.

The revision system

Learning once is not enough. Whatever you learn on a given day, revisit it on day+1, day+3, day+7, and day+14 - not by rereading, but by solving a problem from memory. Day 7, 14, 21, 28, 35, and 42 double as increasingly difficult assessment days layered on top of whatever is scheduled: don't just keep taking in new content on those days.

The blank page test

  • •Spark: write Driver, Executor, Task, Stage, Job, Shuffle, Partition, Broadcast from memory, then explain the whole execution flow.
  • •Any SQL or Python topic: close everything and write the concept's core example from memory before checking it.

Build without a tutorial

  • Airflow: Build a DAG from scratch.
  • Spark: Transform a 5GB dataset.
  • AWS: Design S3 → Glue → Athena.
  • SQL: Solve a business problem.
  • Python: Build an API ingestion script.

Interview depth ladder

Train at all three levels, even though the 45 days target DE1 moving into DE2. Practice drills live in /interview and full design briefs in /system-design.

DE1

  • •Write SQL
  • •Write Python
  • •Build ETL
  • •Use Airflow-shaped orchestration
  • •Use PySpark
  • •Understand cloud storage and IAM
  • •Build data models
  • •Debug a broken pipeline

DE2

  • •Architecture and trade-offs
  • •Performance
  • •Reliability
  • •Data quality
  • •Incremental processing
  • •CDC
  • •Monitoring
  • •Cost
  • •Security
  • •Production debugging

Senior-ish exposure

  • •500 GB/day, 5 TB/day, 500M events/day
  • •99.9% SLA
  • •Late-arriving data
  • •Schema evolution
  • •Data skew
  • •Backfills
  • •Disaster recovery
The interview question bank, by area and rough count
AreaQuestion bank
SQL150 to 200 problems
Python60 to 80 problems
Spark50 to 70 questions
AWS50 questions
Airflow40 questions
Data Engineering100 questions
System Design20 complete designs
Project deep-dives50 questions

Roughly 500+ questions and problems total. The objective is repeated exposure, not memorization.

The Day 45 test

“You're joining an e-commerce company tomorrow. We receive 100M orders/events per day from PostgreSQL, APIs, and Kafka. The analytics team needs customer, order, and revenue datasets by 7 AM every morning. Design and implement the platform.”

Draw the architecture independently, then answer these out loud:

  • •Why Kafka?
  • •Why S3?
  • •Why Parquet?
  • •Why Spark?
  • •Why Redshift?
  • •Why Airflow?
  • •How do you handle duplicates?
  • •How do you handle late data?
  • •How do you implement SCD2?
  • •How do you make it incremental?
  • •How do you handle schema changes?
  • •How do you monitor it?
  • •How do you backfill 6 months?
  • •What happens if Spark fails halfway?
  • •How do you optimize a slow Spark job?
  • •How do you reduce AWS cost?
  • •What if data volume becomes 10x?

If you can answer those and actually implement the core pieces, the 45 days did their job.

What this plan deliberately does not do

It does not spend 45 days teaching Hadoop, Hive, Pig, Scala, Flink, Snowflake, BigQuery, Azure, GCP, Databricks, Kinesis, and the rest, all equally. That produces “I know 20 tools at 20%.” One production stack at 80-90%, with the surrounding ecosystem understood conceptually, is much stronger.

On Lakebench specifically: SQL and Python run for real in the editor. Cloud, orchestration, and streaming are taught as mocked SDKs and brokers in Python. Docker, Kubernetes, Terraform, and a real AWS account are outside this tab, same as they are outside any browser-based course. See the roadmap page for exactly what runs here versus what is homework on your own machine.

Ready to start Day 1?

Core Python and SQL foundational modules are free to explore.

Start Day 1Not sure where to start?

45-Day Plan reviews & ratings

4.9out of 5
840+ student reviews
5 stars
88%
4 stars
9%
3 stars
2%
2 stars
1%
1 star
0%

FAQ

Do I have to finish in exactly 45 days?
No. It is a pace, not a deadline. Falling a week behind just means a 52-day plan, not a failed one. Each phase above still links to the matching Lakebench track so you can pick the pace back up wherever you left off.
Do I really need 8 hours a day?
The 2-hour taught core is the part that has to happen every day. The remaining 6-8 hours compound the material faster, but 3-4 focused hours a day and a longer calendar works too.
Will this get me a data engineer job in 45 days?
No plan can promise that. Finishing every day here means you can write the transforms, the DAG habits, the tests, and the design argument. It does not publish a GitHub repo or sit the interview for you. See the note on experience below.
I already know SQL. Can I skip ahead?
Yes. Start wherever your actual gap is and use this page for pacing the rest.
Does Lakebench run real AWS, Kafka, and Airflow?
No, and it should not pretend to. SQL and Python run for real in the editor. Cloud, orchestration, and streaming are taught as mocked SDKs and brokers in Python. Docker, Kubernetes, Terraform, and a live AWS account are outside this tab. See what maps where on the roadmap page.
LakeBenchPractice today. Build tomorrow.

Warehouse practice that runs in the tab, not on a cluster. Learn concepts, solve interview drills, and mock the round in one place.

Product

  • Studio sandbox
  • Capstone projects

Practice

  • SQL interview questions
  • PySpark interview questions
  • Python interview questions
  • DE theory questions
  • LeetCode for data engineers

Company

  • About
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds & cancellation
  • Shipping & delivery

© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.

No cluster. No install. Just the tab.