Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Learn

Learn data engineering

From complete beginner to job-ready data engineer.

Eleven structured tracks grouped in five phases. Every lesson includes runnable samples and checked exercises in your browser. Mapped lessons hand off to interview drills in Problems. Core foundational modules across every track are free to explore.

Not sure where to begin?

Answer two questions and we'll point you at one first lesson and one free practice ticket.

Where should I start?Jump to Core Python

01Foundations

Python basics and SQL against a warehouse. Start here if you are new.

python

Core Python for Data Engineers

Start from zero: what data engineering is, Python basics, file handling, APIs, and a production ingest capstone. No prior experience needed.

25 lessons · 9 modules · ~284 min

  • Getting Started(2)
  • Foundations(2)
  • Flow, functions & files(5)
  • APIs, streams & objects(3)
  • Time & validation(2)
  • Text & Patterns(2)
  • Reliable Pipelines(4)
  • Packaging & Config(3)
  • Testing & Capstone(2)
Start from scratchRoadmap

sql

SQL & Analytical Warehousing

Start with what data and databases are, then learn SELECT, JOINs, window functions, star schemas, and a fact-table capstone in the SQL editor.

29 lessons · 6 modules · ~332 min

  • Understanding Data(4)
  • Relational core(5)
  • Windows, cleaning & ops(5)
  • Advanced Querying(4)
  • Warehouse & Dimensional Modeling(6)
  • Performance & Production SQL(5)
Start this trackRoadmap

02Data processing

Single-machine pandas, then distributed PySpark when one node is not enough.

pandas

Pandas for Data Manipulation

DataFrames from scratch: reading data, filtering, grouping, merging, reshaping, and a messy-data silver capstone. SQL comparisons at every step.

20 lessons · 5 modules · ~214 min

  • Getting Started with Pandas(4)
  • Frames & cleaning(3)
  • Transforms & shape(6)
  • Scaling Up(3)
  • Production Pandas(4)
Start this trackRoadmap

pyspark

PySpark for Distributed Processing

When data gets too big for one machine: Spark architecture, DataFrames, joins, the medallion pattern, and a bronze-to-gold capstone. Pandas translations included.

21 lessons · 5 modules · ~273 min

  • PySpark Architecture(5)
  • Mental model & DataFrames(3)
  • Aggregations, windows & nested data(6)
  • Performance at Scale(3)
  • Production Spark(4)
Start this trackRoadmap

03Production pipelines

Schedulers, analytics engineering, and quality gates on real queries.

orchestration

Orchestration & Reliable Pipelines

Why pipelines need schedulers, how DAGs work, retries, idempotent loads, backfills, and CDC. Simulated runner (no Airflow scheduler in this tab).

14 lessons · 6 modules · ~180 min

  • What is a Pipeline(3)
  • DAG Mental Model(3)
  • Retries, Idempotency, Dedup(3)
  • Backfills, Catchup, Intervals(2)
  • CDC Concepts(2)
  • Capstone(1)
Start this trackRoadmap

dbt

dbt & Analytics Engineering

What analytics engineering is, how dbt compiles and tests SQL models, staging layers, marts, and incrementals. SQL runs in the editor; no dbt CLI needed.

13 lessons · 6 modules · ~166 min

  • Analytics Engineering(3)
  • What dbt Is(2)
  • Staging(2)
  • unique and relationships(2)
  • Marts and Incremental(3)
  • Capstone(1)
Start this trackRoadmap

quality

Data Quality & Observability

Why data quality matters, how to catch bad data before it reaches dashboards, quality gates, lineage, SLAs, and a publish-suite capstone.

12 lessons · 6 modules · ~148 min

  • What is Data Quality(2)
  • Quality Gates(3)
  • More checks(3)
  • Suites and lineage(2)
  • SLA and SLI(1)
  • Quality capstone(1)
Start this trackRoadmap

04Modern data platforms

Cloud storage and warehouses, streaming concepts, and infra practices.

cloud

Cloud Platforms for Data Engineers

GCP walkthrough with AWS and Azure names: IAM, object storage, BigQuery, Composer, Dataproc, then a free-tier deploy checklist.

14 lessons · 4 modules · ~178 min

  • What is the Cloud(4)
  • Cloud Mental Model(3)
  • Managed Data Services(4)
  • Cost, Security, and Shipping It(3)
Start this trackRoadmap

streaming

Streaming & Message Queues

Why streaming exists, how message queues work, partitions, offsets, consumer groups, late data, and replay. Simulated topics (no broker in this tab).

13 lessons · 6 modules · ~174 min

  • Batch versus Streaming(3)
  • Why Queues(2)
  • Offsets and Groups(2)
  • Delivery and Time(3)
  • Stream-batch(2)
  • Capstone(1)
Start this trackRoadmap

python

Infrastructure & Engineering Practices

Git, Docker, Kubernetes, Terraform, and CI/CD explained from scratch with hands-on critique exercises. The tools run on your laptop; the concepts run here.

10 lessons · 5 modules · ~132 min

  • Git and Linux(3)
  • Docker(2)
  • Kubernetes and Terraform(2)
  • CI/CD(2)
  • Capstone(1)
Start this trackRoadmap

05System design for interviews

Pipeline design tradeoffs and interview-style reasoning, not staff-level architecture.

design

Data Engineering System Design

How to design data systems at scale: batch vs streaming, 500M events/day case studies, cost analysis, and interview preparation. Diagrams and tradeoffs.

9 lessons · 5 modules · ~158 min

  • How to design(2)
  • 500M events/day(3)
  • Serving and cost(2)
  • Interview(1)
  • Capstone(1)
Start this trackRoadmap
View full roadmapOpen problems
LakeBenchPractice today. Build tomorrow.

Warehouse practice that runs in the tab, not on a cluster. Learn concepts, solve interview drills, and mock the round in one place.

Product

  • Studio sandbox
  • Capstone projects

Practice

  • SQL interview questions
  • PySpark interview questions
  • Python interview questions
  • DE theory questions
  • LeetCode for data engineers

Company

  • About
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds & cancellation
  • Shipping & delivery

© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.

No cluster. No install. Just the tab.