
From first SELECT to a design you can defend
Day to day, the job is moving data from a messy source into a shape the business can trust: ingest a file or event stream, put it in a warehouse, model it so analysts do not double-count GMV, schedule that work, test it, and keep it running when a vendor changes a column. The seven phases below map onto eleven Lakebench tracks.
SQL, Python, and PySpark run in browser editors. Cloud, orchestration, and streaming mock external tools in Python. dbt and quality run real queries and frames. System design is diagrams.
Not sure where to begin?
Answer two questions and we'll point you at one first lesson and one free practice ticket.
A working data engineer reads files, writes Python, asks a warehouse questions in SQL, and versions the result in Git. This phase is the one Lakebench can actually run with you in the tab, plus honest Git homework on your laptop.
Milestone
A Python script that reads a CSV, cleans it, and loads it somewhere durable, tracked in Git. Lakebench covers the script. Git and Postgres are homework on your laptop.
Install Git locally and push a private repo
Outside LakebenchVersion the ingest capstone you finish here. Lakebench has no git remote. git add and git commit are taught in the lesson; origin lives on your machine.
Load cleaned CSV into Postgres
Outside LakebenchRun Postgres in Docker or locally. The Core Python ingest capstone produces the clean records; you still have to CREATE TABLE and COPY them yourself.
Indexes, warehouse grain, dimensional models, Pandas on a laptop, and Spark's mental model. Lakebench can teach the operators and the star. It cannot bill a Snowflake account or spin a cluster.
Milestone
A star-schema fact you can defend (grain, additive measures, surrogate keys) plus a PySpark-shaped job on a small frame. The API and the grain rules are the transferable part.
The lesson names Snowflake micro-partitions, BigQuery date partitions, and Redshift SORTKEY. The SQL editor here does not bill bytes the way those warehouses do.
The Python editor here has no pyarrow. You practice iloc chunks, the cousin of chunksize, not an actual Parquet write.
groupBy in this simulator is local pandas. The histogram and the habit of calling explain() are the skill.
A real cloud warehouse account
Outside LakebenchAfter the bytes-scanned lesson in a later phase, open a BigQuery, Snowflake, or Redshift trial and rerun one Lakebench query there. Compare the EXPLAIN. Lakebench will not bill a cloud account for you.
A script you run by hand is not a pipeline. This phase puts retries, edges, backfills, and idempotent loads on a graph. The runner is Python in this tab. Airflow itself is still a laptop install.
Milestone
A DAG-shaped night job: topological order, retries on extract, idempotent load, CDC apply you can replay. Schedule it outside Lakebench.
The lesson is Airflow's mental model in Python dicts. It does not parse a dag.py.
os.environ is the pattern. There is no Secrets Manager in this browser.
Airflow-style DAGs on your laptop
Outside LakebenchInstall Astronomer CLI or apache-airflow locally, or use the free Astro trial. Wrap the ingest capstone as a PythonOperator with retries=3. Lakebench's runner is the mental model, not the scheduler.
Dimensional SQL becomes tested marts. dbt compiles to SQL that really runs here; the CLI is still yours. Quality gates run against the warehouse tables already in this tab. Lineage and SLAs are diagrams until you have a catalog.
Milestone
A staging model, a mart, a unique test that returns zero rows, and a Great-Expectations-shaped suite on df_orders. Then install dbt Core and rebuild it.
A real lineage graph needs OpenLineage, a warehouse INFORMATION_SCHEMA crawl, or dbt docs. The lesson is the shape.
You can compute 19/20 on-time in pandas. You cannot page from an SLO burn-rate alert in this tab.
dbt Core on your laptop
Outside LakebenchInstall dbt Core. Model the star-schema capstone as staging + marts SQL against the SQL editor or your warehouse from phase 5. Lakebench has the compiled SELECT, not a dbt project.
Object storage, IAM, serverless warehouses, and a fake Kafka topic with partitions and late messages. Cloud lessons mock SDKs in Python. Streaming lessons mock a broker in Python. No AWS call and no Kafka process leave this tab.
Milestone
You can name S3/GCS/Blob, price a bytes-scanned query, and reason about offsets, consumer groups, and watermarks. Then open a free-tier project and, separately, a local Kafka or Redpanda if you need the real click.
The PySpark Structured Streaming lesson is the Spark-shaped watermark. This one does not duplicate that exercise.
No live stream in the browser. The lesson teaches event time vs processing time and what a watermark is for.
A real bucket and a real broker
Outside LakebenchCreate one Cloud Storage, S3, or Blob bucket and land the cleaned CSV from phase 1. For Kafka, Redpanda or docker-compose Kafka on a laptop is enough to see a consumer group. Lakebench cannot PUT to a bucket or open a TCP port to a broker.
Docker, Kubernetes, Terraform, and CI/CD do not run in this browser. You will spot bugs in a Dockerfile, name a missing Terraform resource, and order a CI pipeline as Python lists a text checker can grade. The daemon, the cluster, and the apply stay on your machine.
Milestone
Containerize the ingest job, put a CronJob YAML in Git, terraform plan a bucket, and a green CI on main. Lakebench grades the critique. The last mile is on you.
Critique exercise in Python. kind or minikube on a laptop is how you see a pod restart.
Docker Desktop: containerize the ingest job
Outside LakebenchInstall Docker Desktop. Write a Dockerfile around the Python ingest capstone (python:3.12-slim, COPY, USER, CMD). Run it once locally. The lesson taught you not to COPY .env.
kind or minikube: a CronJob that reruns
Outside LakebenchRead a Deployment + CronJob example, then apply it locally. Lakebench is not a cluster.
Terraform CLI: plan a bucket, do not paste secrets
Outside Lakebenchterraform init && terraform plan against a free-tier project. Apply only if you accept the bill. State files stay out of Git.
GitHub Actions (or GitLab CI) on the ingest repo
Outside Lakebenchlint, test, and a dry-run on pull requests. Deploy from main after human review. The ordered list is taught here; the YAML runner is not.
You already built the pieces. This phase is the whiteboard: 500M events/day, batch vs stream, failure, cost, and how to talk it in 45 minutes. No editor. No fake cluster. Then a GitHub README a recruiter can open, and a sober look at certifications.
Milestone
A one-page architecture for a 500M events/day pipeline that names grain, SLA, late data, retry, schema evolution, partitioning, and cost, with a public repo that shows the logic you wrote in Lakebench.
Public GitHub README with grain, schedule, and failure modes
Outside LakebenchRecruiters open GitHub, not a browser sandbox. Copy solutions you wrote, add an architecture diagram, and push. Do not paste secrets.
Badges test product names and exam stamina. They do not replace a repo that states grain. Sit them after you have used the platform, or when a posting you actually want lists the badge. Not before you can write the SQL without the study guide.
| Certification | What it actually tests | Before vs after a first role |
|---|---|---|
| AWS Certified Data Engineer - Associate | Pipelines, data stores, and operations on AWS (Glue, Redshift, Kinesis, Lake Formation, IAM). More product names than grain. | Worth it after you have used S3 and one AWS pipeline at work, or if every posting you want lists AWS. Weak as a substitute for a repo before a first role. |
| Google Professional Data Engineer | BigQuery, Dataflow, Pub/Sub, and designing reliable processing on GCP. Closest to the walkthrough this curriculum uses. | Useful if you already completed a Zoomcamp-style GCP project. Fine after a first role; premature if you have never opened a GCP project. |
| Azure Data Engineer Associate (DP-203) | Azure Data Factory, Synapse, Databricks on Azure, storage, and security in Microsoft shops. | Worth it when your target employers are Azure/enterprise. The exam will not teach you grain. Do the SQL track first. |
| Databricks Certified Data Engineer Associate | Lakehouse, Spark, Delta Lake, Databricks workflows, and Unity Catalog at an associate level. | High signal if you are applying to Databricks-heavy teams. Pair it with the PySpark track plus a real workspace, not this pandas simulator alone. |
| SnowPro Core | Snowflake architecture, warehouses, micro-partitions, Time Travel, RBAC, and SQL on Snowflake. | Reasonable once you have a trial account and can explain clustering vs partitioning. Not a first-job shortcut if you cannot write the SQL without the badge. |
Passing every Lakebench lesson means you can write the transforms, the DAG habits, the tests, and the design argument. It does not publish a repo. The Cloud capstone is a numbered GCP free-tier checklist this tab cannot execute. Do those steps outside, then copy the solutions you wrote, add a README that states grain, schedule, and failure modes, and push that.
Learn capstones