Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Roadmap

From first SELECT to a design you can defend

What a data engineer actually does

Day to day, the job is moving data from a messy source into a shape the business can trust: ingest a file or event stream, put it in a warehouse, model it so analysts do not double-count GMV, schedule that work, test it, and keep it running when a vendor changes a column. The seven phases below map onto eleven Lakebench tracks.

SQL, Python, and PySpark run in browser editors. Cloud, orchestration, and streaming mock external tools in Python. dbt and quality run real queries and frames. System design is diagrams.

Want day-by-day pacing? See the 45-day plan →

Not sure where to begin?

Answer two questions and we'll point you at one first lesson and one free practice ticket.

Where should I start?Browse tracks

01Foundation

A working data engineer reads files, writes Python, asks a warehouse questions in SQL, and versions the result in Git. This phase is the one Lakebench can actually run with you in the tab, plus honest Git homework on your laptop.

Milestone

A Python script that reads a CSV, cleans it, and loads it somewhere durable, tracked in Git. Lakebench covers the script. Git and Postgres are homework on your laptop.

  1. 01
    Python types, control flow, and data structuresLearn here
  2. 02
    Control flow and error handlingLearn here
  3. 03
    Functions, modules, and file I/OLearn here
  4. 04
    REST, generators, and a small connectorLearn here
  5. 05
    Logging, retries, and idempotencyLearn here
  6. 06
    CLI entrypoint and unit testsLearn here
  7. 07
    SQL from SELECT through aggregationsLearn here
  8. 08
    Joins, CTEs, and window functionsLearn here
  9. 09
    Window functionsLearn here
  10. 10
    Git: working tree, index, commitLearn here
  11. 11
    Merge conflicts as an ordered procedureLearn here
  12. 12
    Linux paths, pipes, and exit codesLearn here
  13. 13
    Practice: silver DAG with duplicate grainPractice here
  14. 14

    Install Git locally and push a private repo

    Outside Lakebench

    Version the ingest capstone you finish here. Lakebench has no git remote. git add and git commit are taught in the lesson; origin lives on your machine.

  15. 15

    Load cleaned CSV into Postgres

    Outside Lakebench

    Run Postgres in Docker or locally. The Core Python ingest capstone produces the clean records; you still have to CREATE TABLE and COPY them yourself.

02Data Systems & Modeling

Indexes, warehouse grain, dimensional models, Pandas on a laptop, and Spark's mental model. Lakebench can teach the operators and the star. It cannot bill a Snowflake account or spin a cluster.

Milestone

A star-schema fact you can defend (grain, additive measures, surrogate keys) plus a PySpark-shaped job on a small frame. The API and the grain rules are the transferable part.

  1. 01
    Indexing and query plansLearn here
  2. 02
    Join strategies (hash vs nested loop)Learn here
  3. 03
    OLTP vs OLAP, star schemas, surrogate keysLearn here
  4. 04
    Additive, semi-additive, and non-additive factsLearn here
  5. 05
    SCD Type 1 / 2 / 3 as a taught lessonLearn here
  6. 06
    Warehouse partitioning and clusteringConceptual only, taught here

    The lesson names Snowflake micro-partitions, BigQuery date partitions, and Redshift SORTKEY. The SQL editor here does not bill bytes the way those warehouses do.

  7. 07
    Star-schema fact capstoneLearn here
  8. 08
    Practice: apply city-change CDC (SCD2)Practice here
  9. 09
    Pandas frames, cleaning, and groupbyLearn here
  10. 10
    Parquet and chunked processingConceptual only, taught here

    The Python editor here has no pyarrow. You practice iloc chunks, the cousin of chunksize, not an actual Parquet write.

  11. 11
    PySpark architecture from zeroLearn here
  12. 12
    Spark driver vs executors, lazy plansLearn here
  13. 13
    Joins, broadcast, and partitioningLearn here
  14. 14
    Skew, salting, and Catalyst explainConceptual only, taught here

    groupBy in this simulator is local pandas. The histogram and the habit of calling explain() are the skill.

  15. 15
    Practice: FinOps EXPLAIN budgetPractice here
  16. 16
    Practice: driver OOM from .collect()Practice here
  17. 17

    A real cloud warehouse account

    Outside Lakebench

    After the bytes-scanned lesson in a later phase, open a BigQuery, Snowflake, or Redshift trial and rerun one Lakebench query there. Compare the EXPLAIN. Lakebench will not bill a cloud account for you.

03Orchestration & Reliability

A script you run by hand is not a pipeline. This phase puts retries, edges, backfills, and idempotent loads on a graph. The runner is Python in this tab. Airflow itself is still a laptop install.

Milestone

A DAG-shaped night job: topological order, retries on extract, idempotent load, CDC apply you can replay. Schedule it outside Lakebench.

  1. 01
    Why a scheduler existsLearn here
  2. 02
    Tasks, edges, topological orderLearn here
  3. 03
    Operators, sensors, and XCom (Airflow vocabulary)Learn here

    The lesson is Airflow's mental model in Python dicts. It does not parse a dag.py.

  4. 04
    Retries in the DAG runnerLearn here
  5. 05
    Idempotent loadsLearn here
  6. 06
    Deduplicate before loadLearn here
  7. 07
    Backfills and catchupLearn here
  8. 08
    Data interval vs execution clockLearn here
  9. 09
    CDC concepts and idempotent applyLearn here
  10. 10
    Capstone: a reliable night jobLearn here
  11. 11
    Production loggingLearn here
  12. 12
    Config and secrets (twelve-factor)Conceptual only, taught here

    os.environ is the pattern. There is no Secrets Manager in this browser.

  13. 13
    Practice: broken silver DAG RCAPractice here
  14. 14
    Managed orchestration (Composer / MWAA / Data Factory)Learn here
  15. 15

    Airflow-style DAGs on your laptop

    Outside Lakebench

    Install Astronomer CLI or apache-airflow locally, or use the free Astro trial. Wrap the ingest capstone as a PythonOperator with retries=3. Lakebench's runner is the mental model, not the scheduler.

04Analytics Engineering & Quality

Dimensional SQL becomes tested marts. dbt compiles to SQL that really runs here; the CLI is still yours. Quality gates run against the warehouse tables already in this tab. Lineage and SLAs are diagrams until you have a catalog.

Milestone

A staging model, a mart, a unique test that returns zero rows, and a Great-Expectations-shaped suite on df_orders. Then install dbt Core and rebuild it.

  1. 01
    What dbt is (SQL runs, CLI does not)Learn here
  2. 02
    source() and ref() as compiled namesLearn here
  3. 03
    Staging modelsLearn here
  4. 04
    not_null, unique, and relationships testsLearn here
  5. 05
    Marts from staging CTEsLearn here
  6. 06
    Incremental vs full refreshLearn here
  7. 07
    dbt capstone: staging, mart, testLearn here
  8. 08
    SQL constraints and data qualityLearn here
  9. 09
    Pandas schema contractsLearn here
  10. 10
    Why quality gatesLearn here
  11. 11
    expect_column_not_null / unique / row_countLearn here
  12. 12
    Run a suite and fail the jobLearn here
  13. 13
    Data lineage (diagram, not a catalog)Conceptual only, taught here

    A real lineage graph needs OpenLineage, a warehouse INFORMATION_SCHEMA crawl, or dbt docs. The lesson is the shape.

  14. 14
    SLA / SLI (diagram plus a toy ratio)Conceptual only, taught here

    You can compute 19/20 on-time in pandas. You cannot page from an SLO burn-rate alert in this tab.

  15. 15
    Quality suite capstoneLearn here
  16. 16
    Pandas messy-data silver capstoneLearn here
  17. 17

    dbt Core on your laptop

    Outside Lakebench

    Install dbt Core. Model the star-schema capstone as staging + marts SQL against the SQL editor or your warehouse from phase 5. Lakebench has the compiled SELECT, not a dbt project.

05Cloud & Streaming

Object storage, IAM, serverless warehouses, and a fake Kafka topic with partitions and late messages. Cloud lessons mock SDKs in Python. Streaming lessons mock a broker in Python. No AWS call and no Kafka process leave this tab.

Milestone

You can name S3/GCS/Blob, price a bytes-scanned query, and reason about offsets, consumer groups, and watermarks. Then open a free-tier project and, separately, a local Kafka or Redpanda if you need the real click.

  1. 01
    Why the cloud, and which oneLearn here
  2. 02
    IAM: who can do whatLearn here
  3. 03
    Object storage as the bronze landing zoneLearn here
  4. 04
    Serverless SQL warehouses (bytes scanned)Learn here
  5. 05
    Managed SparkLearn here
  6. 06
    Serverless compute (functions)Learn here
  7. 07
    Why a message queueLearn here
  8. 08
    Topics, partitions, and keysLearn here
  9. 09
    Offsets, commits, consumer groupsLearn here
  10. 10
    At-least-once vs exactly-onceLearn here
  11. 11
    Late and out-of-order messagesLearn here
  12. 12
    Watermarks (Kafka-shaped cousin of Spark)Learn here

    The PySpark Structured Streaming lesson is the Spark-shaped watermark. This one does not duplicate that exercise.

  13. 13
    Structured Streaming and watermarks (PySpark)Conceptual only, taught here

    No live stream in the browser. The lesson teaches event time vs processing time and what a watermark is for.

  14. 14
    Stream-batch unificationLearn here
  15. 15
    Streaming capstone: late messagesLearn here
  16. 16
    PySpark medallion, then incremental MERGELearn here
  17. 17

    A real bucket and a real broker

    Outside Lakebench

    Create one Cloud Storage, S3, or Blob bucket and land the cleaned CSV from phase 1. For Kafka, Redpanda or docker-compose Kafka on a laptop is enough to see a consumer group. Lakebench cannot PUT to a bucket or open a TCP port to a broker.

06Infrastructure & Production Engineering

Docker, Kubernetes, Terraform, and CI/CD do not run in this browser. You will spot bugs in a Dockerfile, name a missing Terraform resource, and order a CI pipeline as Python lists a text checker can grade. The daemon, the cluster, and the apply stay on your machine.

Milestone

Containerize the ingest job, put a CronJob YAML in Git, terraform plan a bucket, and a green CI on main. Lakebench grades the critique. The last mile is on you.

  1. 01
    Read a Dockerfile (spot the .env copy)Learn here
  2. 02
    Dockerfile least-privilege USERLearn here
  3. 03
    Kubernetes CronJob vs DeploymentLearn here

    Critique exercise in Python. kind or minikube on a laptop is how you see a pod restart.

  4. 04
    Terraform: which resource is missingLearn here
  5. 05
    CI/CD stages: lint, test, build, deployLearn here
  6. 06
    Order the CI steps on mainLearn here
  7. 07
    Ship checklist capstoneLearn here
  8. 08
    Reading a cloud billLearn here
  9. 09
    VPC and private connectivityLearn here
  10. 10
    Capstone: deploy ingest to the cloud (checklist)Learn here
  11. 11

    Docker Desktop: containerize the ingest job

    Outside Lakebench

    Install Docker Desktop. Write a Dockerfile around the Python ingest capstone (python:3.12-slim, COPY, USER, CMD). Run it once locally. The lesson taught you not to COPY .env.

  12. 12

    kind or minikube: a CronJob that reruns

    Outside Lakebench

    Read a Deployment + CronJob example, then apply it locally. Lakebench is not a cluster.

  13. 13

    Terraform CLI: plan a bucket, do not paste secrets

    Outside Lakebench

    terraform init && terraform plan against a free-tier project. Apply only if you accept the bill. State files stay out of Git.

  14. 14

    GitHub Actions (or GitLab CI) on the ingest repo

    Outside Lakebench

    lint, test, and a dry-run on pull requests. Deploy from main after human review. The ordered list is taught here; the YAML runner is not.

07System Design & Career

You already built the pieces. This phase is the whiteboard: 500M events/day, batch vs stream, failure, cost, and how to talk it in 45 minutes. No editor. No fake cluster. Then a GitHub README a recruiter can open, and a sober look at certifications.

Milestone

A one-page architecture for a 500M events/day pipeline that names grain, SLA, late data, retry, schema evolution, partitioning, and cost, with a public repo that shows the logic you wrote in Lakebench.

  1. 01
    Read the brief: volume, SLA, late data, consumersConceptual only, taught here
  2. 02
    Batch vs stream as a decision, not a fashionConceptual only, taught here
  3. 03
    500M events/day: ingestConceptual only, taught here
  4. 04
    500M events/day: storage, partitions, lakehouse formatsConceptual only, taught here
  5. 05
    500M events/day: failure, retry, schema evolutionConceptual only, taught here
  6. 06
    Serving analyticsConceptual only, taught here
  7. 07
    Cost and FinOps in a design reviewConceptual only, taught here
  8. 08
    The 45-minute design interviewConceptual only, taught here
  9. 09
    Capstone: full pipeline designConceptual only, taught here
  10. 10
    Core Python ingest capstone (logic you already wrote)Learn here
  11. 11

    Public GitHub README with grain, schedule, and failure modes

    Outside Lakebench

    Recruiters open GitHub, not a browser sandbox. Copy solutions you wrote, add an architecture diagram, and push. Do not paste secrets.

Certifications

Badges test product names and exam stamina. They do not replace a repo that states grain. Sit them after you have used the platform, or when a posting you actually want lists the badge. Not before you can write the SQL without the study guide.

Data engineering certifications, what they test, and when they are worth it
CertificationWhat it actually testsBefore vs after a first role
AWS Certified Data Engineer - AssociatePipelines, data stores, and operations on AWS (Glue, Redshift, Kinesis, Lake Formation, IAM). More product names than grain.Worth it after you have used S3 and one AWS pipeline at work, or if every posting you want lists AWS. Weak as a substitute for a repo before a first role.
Google Professional Data EngineerBigQuery, Dataflow, Pub/Sub, and designing reliable processing on GCP. Closest to the walkthrough this curriculum uses.Useful if you already completed a Zoomcamp-style GCP project. Fine after a first role; premature if you have never opened a GCP project.
Azure Data Engineer Associate (DP-203)Azure Data Factory, Synapse, Databricks on Azure, storage, and security in Microsoft shops.Worth it when your target employers are Azure/enterprise. The exam will not teach you grain. Do the SQL track first.
Databricks Certified Data Engineer AssociateLakehouse, Spark, Delta Lake, Databricks workflows, and Unity Catalog at an associate level.High signal if you are applying to Databricks-heavy teams. Pair it with the PySpark track plus a real workspace, not this pandas simulator alone.
SnowPro CoreSnowflake architecture, warehouses, micro-partitions, Time Travel, RBAC, and SQL on Snowflake.Reasonable once you have a trial account and can explain clustering vs partitioning. Not a first-job shortcut if you cannot write the SQL without the badge.

Capstone: logic here, artifact on GitHub

Passing every Lakebench lesson means you can write the transforms, the DAG habits, the tests, and the design argument. It does not publish a repo. The Cloud capstone is a numbered GCP free-tier checklist this tab cannot execute. Do those steps outside, then copy the solutions you wrote, add a README that states grain, schedule, and failure modes, and push that.

All capstones on one page

Learn capstones

  • ingest-capstone
  • star-schema-capstone
  • messy-data-capstone
  • medallion-incremental
  • medallion-pipeline
  • deploy-ingest-capstone
  • reliable-night-job
  • staging-mart-test
  • late-messages-capstone
  • quality-suite-capstone
  • ship-checklist
  • full-pipeline-design

Studio interview drills

  • Broken DAG: silver_events unique test red
  • FinOps: dashboard GMV is scanning the lake
  • SCD Type 2: apply city-change CDC
  • Broken job: driver OOM from .collect()
LakeBenchPractice today. Build tomorrow.

Warehouse practice that runs in the tab, not on a cluster. Learn concepts, solve interview drills, and mock the round in one place.

Product

  • Studio sandbox
  • Capstone projects

Practice

  • SQL interview questions
  • PySpark interview questions
  • Python interview questions
  • DE theory questions
  • LeetCode for data engineers

Company

  • About
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds & cancellation
  • Shipping & delivery

© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.

No cluster. No install. Just the tab.