
45-day sprint
Not “learn data engineering.” The goal is narrower: can you sit in a DE1 or DE2 interview, take a real pipeline problem, design it, code it, explain the trade-offs, debug it, and defend your decisions. Two hours taught a day, plus deliberate practice, across 11 phases and 45 days. One coherent stack instead of twenty tools at twenty percent each.
The Day 45 Benchmark
By Day 45 you should be able to walk through this lifecycle out loud without hesitation: defending table grain, partition keys, compute trade-offs, and failure recovery across every link in the chain. Click any skill to jump directly to its lesson:
Not sure where to begin?
Answer two questions and we'll point you at one first lesson and one free practice ticket.
1 curriculum day daily (4 to 8 hours deliberate practice)
2 hours taught (concept, live coding, interview questions, build something), then 6-8 hours of deliberate practice. Not “go watch 5 hours of YouTube.” Every day ends with a specific engineering assignment.
| Block | Time | What you do |
|---|---|---|
| Concept | 30m | The idea, in plain language, before any syntax. |
| Live coding / architecture | 30m | Watch or work through the pattern being built, not just described. |
| Interview questions | 30m | The questions this topic actually gets asked in interviews. |
| Build something | 30m | A small, real example of the day's concept, not a toy snippet. |
| Rebuild without looking | 1h | Close everything. Rebuild the day's core example from memory, then compare. |
| Project implementation | 2h | Apply today's concept to whichever capstone project you picked. |
| SQL/Python practice | 1h | Drills unrelated to today's topic, to keep older material warm. |
| Interview questions | 1h | Answer today's prompts out loud, not just in your head. |
| Revision | 30m | Redo one exercise from 1, 3, 7, and 14 days ago from memory. |
| Engineering journal | 30m | What you learned, what you implemented, what broke, why it broke, how you fixed it, what trade-off you made, and which interview questions you couldn't answer. |
That is roughly 8 hours a day. The final 30 minutes, the journal, is the one that turns Day 45 into a personal knowledge base.
Eleven phases, one coherent stack instead of twenty tools at twenty percent. Each phase links to the closest matching Lakebench track, but the day list below is the exact plan, not a re-sorted version of it.
SQL is the single skill worth over-investing in: not basic SELECT queries, but difficult interview problems solved without panic.
Teaches: SELECT, WHERE, ORDER BY, DISTINCT, LIMIT, aliases, NULL, CASE, CAST, basic functions, why the database executes clauses in a fixed logical order
Build: customers, orders, products, and payments tables. Solve 30 queries against them.
Teaches: GROUP BY, HAVING, COUNT, COUNT DISTINCT, SUM, AVG, MIN/MAX
Teaches: INNER, LEFT, RIGHT, FULL, CROSS, SELF JOIN
Teaches: ROW_NUMBER, RANK, DENSE_RANK, LAG, LEAD, SUM OVER, AVG OVER, PARTITION BY, ORDER BY
Teaches: CTEs, recursive CTE concept, subqueries, correlated subqueries, EXISTS, UNION, UNION ALL, INTERSECT, EXCEPT
Teaches: indexes, query plans, EXPLAIN, EXPLAIN ANALYZE, partitioning, clustering concepts, predicate pushdown, avoiding SELECT *, joins and cardinality, materialized views
Build: Deliberately write a slow query, then optimize it.
No new concepts. A 3-hour mock SQL exam: 10 easy, 10 medium, 10 hard, and 5 real-world DE problems. Then 2 hours reviewing every mistake. By the end of today, SQL should be interview-ready.
Not turning you into a software engineer. The question is narrower: can you write production-quality Python that processes data and builds pipelines?
Teaches: data types, lists & dicts, hash sets, loops, functions, comprehensions, csv.DictReader
Build: Dirty CSV ingestion → parsing & validation → clean CSV export.
Data engineering Python: safe string parsing, currency precision, O(1) set lookups, and pure functions.
Teaches: modules, exception handling, logging levels, file descriptors, JSON & CSV, env vars
Build: API → Python → JSON validation → CSV export with structured logging.
Production habits: targeted exceptions, log levels, file handling, and 12-factor configuration.
Teaches: read_csv, read_json, merge, groupby, vectorization, missing values, datetime, aggregation
Build: orders.csv + customers.json → Pandas clean & merge → aggregate report + data quality summary.
Know Pandas for local profiling and wrangling, but prioritize vectorization over slow apply loops.
Teaches: config management, structured logging, custom exceptions, retries & backoff, idempotency, pytest
Build: Modular pipeline (extract, transform, load, config, utils) with idempotent upserts and automated tests.
Production pipelines survive network blips, malformed payloads, and repeated Airflow retries safely.
Teaches: REST extraction, schema validation, dead-letter quarantine, pure transforms, idempotent upsert, pytest
Build: REST API → validation & DLQ quarantine → transform → database upsert with failure injection.
Phase 2 Capstone: validate schemas, isolate bad rows in quarantine, handle 503 retries, and guarantee idempotency.
Now you learn what data engineering actually is, not just the tools around it.
Teaches: ETL vs ELT, batch vs streaming, micro-batching, medallion stages, raw immutable storage, replay recovery
Build: Multi-source ingestion → immutable partitioned raw storage → dual ETL vs ELT comparison → analytics datamart.
System trade-offs: compute costs of pre-load transforms vs in-warehouse SQL, and raw storage as replay insurance.
Teaches: OLTP vs OLAP, fact tables, dimension tables, star vs snowflake, surrogate vs natural keys, declaring grain
Build: Kimball Star Schema in SQLite (orders, customers, products, dates) with grain fan-out debugger.
Declare the grain in plain English first: grain dictates primary keys, valid metrics, and prevents join fan-out.
Teaches: SCD Type 0 (Freeze), SCD Type 1 (Replace), SCD Type 2 (Version), valid_from & valid_to, is_current flag, point-in-time joins
Build: Multi-batch SCD Type 2 dimension engine with surrogate keys, half-open intervals, and time-travel SQL.
Mental model: Freeze (0), Replace (1), Version (2). Use half-open intervals [start, end) for point-in-time fact joins.
Teaches: lake vs warehouse vs lakehouse, Medallion architecture, Parquet vs CSV, Hive partitioning, partition pruning, small files compaction
Build: Medallion Data Lake engine: raw CSV → bronze standardization → silver partitioned Parquet → gold daily sales aggregate.
Never destroy raw data. Store columnar Parquet with coarse temporal partitions so engines skip 90%+ of files.
Teaches: multi-source integration, read-replica CDC, API exponential backoff & DLQ, Kafka stream dedup, dual-speed SLAs, pipeline idempotency
Build: Enterprise multi-source platform: PostgreSQL + flaky REST API (with DLQ) + Kafka clickstream → S3 Parquet lake + SCD2 warehouse.
Phase 3 Capstone: design first, hit architectural failure modes, correct the design, and defend trade-offs.
Cloud, now that the fundamentals are in place. Same 5 days, same shape - pick AWS, GCP, or Azure below.
Teaches: IAM roles & policies, STS temporary tokens, least privilege, S3 object model, EC2 vs Lambda, CloudWatch alarms
Build: AWS security simulator: IAM evaluation, temporary STS token assumption, and CloudWatch metric alerting.
Never hardcode AWS keys. Attach IAM roles to compute and issue short-lived temporary tokens via STS.
Teaches: S3 prefixes vs directories, Medallion zones (Raw/Silver/Gold), Hive partitioning, partition pruning, lifecycle tiering, compaction
Build: S3 Medallion Lake engine: raw ingestion → schema validation → Hive partitioned Parquet → curated sales mart.
Partition by query access patterns, keep immutable raw data for replay, and automate lifecycle tiering to Glacier.
Teaches: S3 + Glue + Athena flow, Glue Data Catalog, external tables & SerDes, Athena scan pricing ($5/TB), partition pruning, schema evolution
Build: Serverless query engine: Glue catalog discovery → partition pruning benchmarks → pre-warehouse data quality validation.
Athena is serverless compute charging by data scanned. Pair Hive-partitioned Parquet with column pruning to cut 99% of costs.
Teaches: Redshift MPP architecture, DISTSTYLE (KEY, ALL, EVEN), Sort Keys & Zone Maps, columnar compression, COPY from S3, atomic staging merges
Build: Redshift MPP warehouse engine: parallel S3 COPY → staging deduplication → atomic idempotent merge transaction.
Redshift doesn't enforce primary key constraints on write. Guarantee idempotency using staging tables and atomic transactions.
Teaches: complete AWS platform flow, high-watermark extraction, S3 lakehouse staging, Glue catalog registration, Redshift COPY & merge, reconciliation
Build: End-to-end AWS data platform: PostgreSQL incremental extraction → S3 raw/curated → Athena validation → Redshift merge.
Decouple Postgres from Redshift using S3. Advance watermarks only after warehouse commits, and use LEFT JOIN fallbacks.
This is where you become much more employable. Current DE-II postings repeatedly emphasize Spark and scalable pipeline processing.
Teaches: horizontal scaling, driver vs executor, Catalyst DAG, transformations vs actions, narrow vs wide stages, fault tolerance lineage
Build: Distributed execution simulator: driver scheduling, narrow transformations, shuffle partitions, and worker failure recovery.
Spark coordinates distributed compute. Never run df.collect() on large data, and remember files in S3 are not partitions.
Teaches: lazy evaluation, the 9 core DataFrame ops, predicate pushdown, bitwise filtering (&, |), avoiding withColumn loops, window deduplication
Build: E-commerce PySpark pipeline: raw CSV ingestion → data quality filtering → multi-metric aggregation → master data join.
Never use Python 'and'/'or' on Spark Columns: use '&' and '|'. Avoid chaining withColumn in loops or using df.distinct() blindly.
Teaches: partition boundaries, narrow vs wide transformations, shuffle write & fetch, repartition vs coalesce, coalesce(1) single-core hazard, partition sizing
Build: Multi-stage pipeline: narrow filtering → wide shuffle aggregation → hash repartitioning by customer key → coalesce(2) lake write.
Never use coalesce(1) before writing large datasets: it collapses upstream execution to a single core and crashes workers.
Teaches: BroadcastHashJoin (<100MB), partition pruning, Parquet predicate pushdown, when to cache vs unpersist, data skew key salting, small files coalesce
Build: Refactor a slow Spark job with 6 bottlenecks (wide scan, missing broadcast, shuffle skew, coalesce(1)) into a 95% faster pipeline.
The fastest data is data Spark never reads or moves. Use broadcast joins for small dimensions and salt skewed keys.
Teaches: data contracts & early validation, string & enum sanitation, windowed latest-record deduplication, join cardinality diagnostics, pre-write circuit breakers, idempotent writes
Build: End-to-end PySpark pipeline: raw ingestion → contract validation → window deduplication → broadcast join → pre-write assertions.
Verify dimension uniqueness before joining to prevent row explosion, and write with mode('overwrite') to ensure rerun safety.
Teaches: Spark execution mental model, Spark UI triage framework, shuffle mechanics & skew, broadcast trade-offs, repartition vs coalesce, executor failure triage
Build: Spark Interview Lab: 5 TB skew autopsy with salting, shuffle diagnostic analyzer, and small files compaction optimizer.
Never just say 'add more executors' when a job is slow. Triage the Spark UI: find the slow stage, inspect task skew, and check shuffle volume.
Orchestration coordinates the moving parts of modern data systems. Build production DAGs, handle backfills safely, prevent worker starvation, and master failure triage.
Teaches: DAGs as directed acyclic graphs, Scheduler vs Webserver vs Executor vs Worker, Task state lifecycle & cascades, Operators vs Sensors vs TaskFlow API, Bitshift dependency chaining (>>), XCom mechanics & metadata DB hazards
Build: Airflow execution engine: DAG registry, topological sort dependency resolution, state transition simulation, and 48KB XCom overflow protection.
Never pass dataframes or raw datasets through XCom. Use XCom strictly for lightweight metadata and remote pointers (S3/GCS paths).
Teaches: Logical date vs Execution date vs Data interval, catchup=False vs catchup=True, CLI backfills & idempotent rerun safety, Sensor poke vs reschedule mode, Worker slot starvation prevention, Retries with exponential backoff & jitter, Avoiding top-level Variable.get() / Connection.get() DB blasts
Build: Production Airflow simulator: logical date intervals, sensor slot starvation remediation, exponential retry jitter, and execution_timeout circuit breakers.
Never use sensor poke mode for long waits (>5 min): it locks worker slots and starves all other pipelines. Always use mode='reschedule'.
Teaches: Separation of concerns: Airflow orchestrates, Spark computes, Why running PySpark code inside Airflow workers is a critical anti-pattern, SparkSubmitOperator vs EMR / Dataproc / Databricks operators, Deferrable operators & asynchronous triggers, S3/GCS staging handoff: Extract → Raw → Spark → Curated → Warehouse COPY/MERGE, Atomic partition replacement & staging cleanup
Build: Cloud lakehouse orchestrator: Airflow async trigger DAG → Dataproc/EMR Spark execution → S3 Parquet validation → atomic warehouse MERGE gate.
Never execute heavy compute on Airflow worker nodes. Treat Airflow strictly as the air traffic controller dispatching jobs to managed clusters.
Teaches: The 6 classic production orchestration failures, API rate limits (429) & exponential jitter backoff, Cloud storage eventual consistency & missing partition races, Source schema drift & dead-letter quarantine, Database connection pool exhaustion & connection pooling, Spark executor OOM retry cascade prevention, Metadata DB lock contention triage
Build: Production Incident Triage Console: live failure injection (429 rate limit, schema drift, DB pool exhaustion, storage lag) with diagnostic logs, real-time remediation, and health verification.
When a pipeline fails, never blindly click 'Clear' to retry. First identify whether the failure is transient, systemic, or a data-corruption hazard.
Analytics engineering turns raw data lakes into trusted, tested data models. Master SQL-first transformations, modular DAG refs, data contracts, and watermarked incremental processing.
Teaches: dbt architecture & compilation to native SQL, Sources, staging models & column renaming, ref() DAG generation & lineage, Generic tests (not_null, unique, accepted_values, relationships), dbt macros & Jinja templating, SCD Type 2 snapshots (check vs timestamp strategy)
Build: dbt transformation engine: raw sources → staging clean views → dimensional fact mart with automated lineage graph compilation and snapshot history tracking.
Never hardcode database or schema names in SQL. Always use {{ source() }} and {{ ref() }} to guarantee atomic DAG compilation and environment isolation.
Teaches: Defense in depth: Pre-load vs In-flight vs Post-load verification, Column invariants (uniqueness, referential integrity, range bounds), Row count anomaly detection & statistical deviation (Z-score), Freshness SLAs & silent pipeline stalls, Circuit breakers: quarantine vs pipeline abort, Data contract enforcement
Build: Data quality & circuit breaker engine: schema enforcement → statistical volume anomaly detector → referential integrity validator → quarantine isolation router.
A failing pipeline that halts execution is an inconvenience; a silent corrupt pipeline that writes dirty metrics to production is an executive catastrophe.
Teaches: Full refresh vs incremental processing trade-offs, Watermark tracking & high-watermark state stores, Lookback windows for late-arriving records, Atomic MERGE & upsert mechanics (is_incremental() macro), Unique key deduplication before warehouse write, Handling hard deletes and tombstones
Build: Incremental lakehouse pipeline: high-watermark state tracking → late-arriving event lookback buffer → idempotent atomic MERGE → audit log reconciliation.
Always include a lookback window (e.g. current_watermark - interval '3 hours') in incremental filters to catch out-of-order and late-arriving records.
A complete, production-grade batch data platform: PostgreSQL extraction → Airflow orchestration → S3 Bronze staging → PySpark Silver cleaning & Gold aggregation → Redshift/BigQuery warehouse → dbt metrics → SLA alerting.
Teaches: Transactional database CDC vs batch extraction, High-watermark state persistence in DynamoDB/PostgreSQL, Partitioned S3/GCS Bronze raw lake staging, Schema validation & payload checksum hashing, REST API event enrichment ingestion, Idempotent extraction chunking
Build: Platform Stage 1: PostgreSQL & Stripe API batch extractor → checksum validator → partitioned Hive Parquet S3 Bronze lake writer.
Never perform transformations during the raw extraction phase. Write Bronze data in its raw, unmodified form to enable historical replay.
Teaches: Medallion curation: Bronze raw to Silver cleaned to Gold aggregates, Windowed deduplication by primary key and updated_at timestamp, Data contract casting & null substitution policies, Broadcast dimension enrichment, Gold daily revenue & customer lifetime value fact tables, Atomic S3 partition swaps
Build: Platform Stage 2: PySpark Silver deduplication & validation engine → Gold metric aggregations → broadcast join enricher → atomic lakehouse writer.
Always deduplicate Silver data using row_number() over (partition by id order by updated_at desc) before performing any join operations.
Teaches: End-to-end Airflow DAG tying extraction, Spark, and dbt together, Warehouse atomic MERGE / COPY ingestion gates, dbt data tests as blocking deployment gates, SLA monitoring & PagerDuty alert routing, Complete system architecture verbal defense, Cost & performance trade-off justification
Build: Platform Stage 3: End-to-end Airflow master DAG orchestrating extraction, EMR Spark, Redshift COPY, dbt tests, and Slack/PagerDuty webhook alerts.
In a Senior DE interview, walk through this platform using the 'Requirements → Bottlenecks → Trade-offs → Failure Modes' framework.
Modern data platforms cannot afford 24-hour batch latency. Master real-time change data capture, distributed log queues, event ordering, and streaming state engines.
Teaches: Distributed commit log architecture, Brokers, topics, partitions & segments, Message keys & deterministic partition routing, Producers: acks (0, 1, all) & idempotence, Consumer groups, rebalancing & offset commit semantics, Lag monitoring & consumer scale limits
Build: Kafka streaming simulator: producer partition router with hashing, multi-consumer group coordinator, offset commit manager, and lag spike monitor.
Partition count is the fundamental unit of parallelism. You cannot have more active consumers in a consumer group than partitions in a topic.
Teaches: Change Data Capture (CDC) via database write-ahead log (WAL / binlog), Debezium connector architecture & event envelopes (before, after, op), Out-of-order events & watermark event-time processing, Deduplication & exactly-once processing guarantees, Streaming stateful joins (stream-stream vs stream-table), Changelog table compaction
Build: CDC stream processing engine: PostgreSQL WAL log parser → Debezium event envelope flattener → out-of-order watermark buffer → compacted lakehouse sink.
Never query source databases directly for real-time changes. Read the database WAL via Debezium to avoid impacting transactional query performance.
What separates a junior coder from a Senior Staff Data Engineer: resilient distributed systems design, airtight production operations, CI/CD, and live incident triage.
Teaches: 500M events/day data platform blueprint, Throughput, storage & compute capacity sizing, Lambda vs Kappa vs Delta lakehouse architectures, Decoupling ingestion, processing, and analytical serving, Multi-tenant isolation & noisy neighbor prevention, Scaling from 500M to 5B events/day: 10x bottlenecks
Build: System Design interactive sizing calculator: daily event volume → throughput/sec → storage decay → worker cluster sizing → cost estimation model.
In system design interviews, always start with requirements, SLA numbers, and back-of-the-envelope calculations before drawing single boxes or choosing technologies.
Teaches: Git trunk-based development & feature branching, Data pipeline CI/CD with automated testing stages, Docker containerization for reproducible Spark/Airflow environments, Secrets management & credential rotation (IAM, Vault, AWS Secrets Manager), Infrastructure as Code (Terraform) principles, Observability: Metrics, logs, traces, and SLO/SLA error budgets
Build: Production CI/CD & validation pipeline: GitHub Actions simulator → linting → SQL fluff → PySpark unit tests → Docker build → staging deployment gate.
Never test data pipelines in production. Enforce automated pre-merge testing with mock datasets to catch bugs before they corrupt production tables.
Teaches: The 6 mission-critical production incident drills, Scenario 1: Pipeline takes 3 hours instead of 20 mins, Scenario 2: Spark job sudden executor OOM crash, Scenario 3: Financial revenue metric jumps 30% unexpectedly, Scenario 4: Pipeline reports success but destination table is empty, Scenario 5: Duplicate records appear in downstream mart, Scenario 6: Upstream source introduces unannounced schema drift
Build: Production Incident War Room: real-time incident simulator where student investigates symptoms, queries telemetry, diagnoses root cause, and applies permanent fix.
Treat every incident as a systemic failure, not human error. Document the root cause timeline and implement automated regression tests to prevent recurrence.
The culmination of your 45-day journey. Build an enterprise-grade real-time streaming platform from scratch, then defend your knowledge across 7 rigorous mock interview rounds.
Teaches: Real-time e-commerce architecture specification, PostgreSQL transactional source → Debezium CDC → Kafka event bus, PySpark Structured Streaming with watermarks & stateful deduplication, S3/ADLS Medallion lakehouse (Delta/Iceberg/Parquet), Warehouse dimension modeling & incremental dbt mart, Production monitoring, alerting & end-to-end data contracts
Build: Capstone Platform: PostgreSQL CDC → Kafka streaming pipeline → PySpark streaming enrichment → Medallion lakehouse → dbt analytical warehouse.
Build every layer independently and verify contracts at each boundary. If a bug occurs, isolate whether the source is CDC, Kafka, Spark, or the warehouse.
Teaches: The 7 rounds of senior data engineering hiring, Round 1: Advanced SQL & Data Modeling (60 min), Round 2: Python Data Structures & Core Algorithms (45 min), Round 3: PySpark & Distributed Computing Deep Dive (45 min), Round 4: Data Engineering Pipeline Architecture & Storage (60 min), Round 5: Distributed System Design at 500M+ Scale (60 min), Round 6: Project Portfolio Deep Dive & Battle Scars (60 min), Round 7: Senior Staff Behavioral & Leadership Defense (30 min)
Build: Full 7-Round Interview Simulator: interactive question prompt, rubric evaluation, weak answer red flags, model senior staff responses, and scoring scorecard.
Senior engineers don't win interviews by reciting syntax. They win by explaining trade-offs, anticipating edge cases, and showing proven production empathy.
Pick one domain and build all three against it so the story stays coherent end to end. These map onto the same capstones you already write inside each Learn track, described the way an interviewer would ask about them. The full list, with links to every capstone lesson, lives on /projects.
The primary batch project. Everything else references it.
Sources: PostgreSQL, REST API, CSV. Features: incremental loads, SCD2, data quality, partitioning, deduplication, schema validation, logging, monitoring.
Tables: customers, products, orders, order_items, payments, shipments, returns, reviews
Same e-commerce domain as A, so the two connect. The streaming project.
Application → Kafka → Spark Streaming → S3 → Redshift. Teaches partitions, offsets, consumer groups, event ordering, late events, duplicates, and watermarks.
Tables: page_view, product_view, add_to_cart, checkout, purchase
The queue is taught and practiced as a mocked broker in Python, per the Streaming track. No Kafka process runs in this tab.
The project for the strongest DE2-depth conversations.
Operational DB → CDC → Kafka → Spark → S3 → Warehouse → dbt → Customer 360. Implements SCD2, CDC, MERGE, incremental processing, data quality, and schema evolution.
Tables: customer, customer_address, customer_orders, customer_payments, customer_interactions
Do not claim you built these systems at an employer when you did not. Interviewers drill: how many records, what was your SLA, what incident happened, who were your consumers, what was your AWS bill, how did you handle on-call. A fabricated story eventually collapses.
Instead, build the three projects above to a level where they genuinely resemble production systems, put them in a public GitHub repo with an architecture diagram, README, sample data, tests, and documented trade-offs, and describe them accurately as hands-on capstone work. If you already worked at a company for a year or two, the stronger move is to retrofit what you genuinely learned there onto the real business context you already know, so a 45-minute deep dive draws on things you actually understand rather than a memorized lie.
Learning once is not enough. Whatever you learn on a given day, revisit it on day+1, day+3, day+7, and day+14 - not by rereading, but by solving a problem from memory. Day 7, 14, 21, 28, 35, and 42 double as increasingly difficult assessment days layered on top of whatever is scheduled: don't just keep taking in new content on those days.
Train at all three levels, even though the 45 days target DE1 moving into DE2. Practice drills live in /interview and full design briefs in /system-design.
| Area | Question bank |
|---|---|
| SQL | 150 to 200 problems |
| Python | 60 to 80 problems |
| Spark | 50 to 70 questions |
| AWS | 50 questions |
| Airflow | 40 questions |
| Data Engineering | 100 questions |
| System Design | 20 complete designs |
| Project deep-dives | 50 questions |
Roughly 500+ questions and problems total. The objective is repeated exposure, not memorization.
“You're joining an e-commerce company tomorrow. We receive 100M orders/events per day from PostgreSQL, APIs, and Kafka. The analytics team needs customer, order, and revenue datasets by 7 AM every morning. Design and implement the platform.”
Draw the architecture independently, then answer these out loud:
If you can answer those and actually implement the core pieces, the 45 days did their job.
It does not spend 45 days teaching Hadoop, Hive, Pig, Scala, Flink, Snowflake, BigQuery, Azure, GCP, Databricks, Kinesis, and the rest, all equally. That produces “I know 20 tools at 20%.” One production stack at 80-90%, with the surrounding ecosystem understood conceptually, is much stronger.
On Lakebench specifically: SQL and Python run for real in the editor. Cloud, orchestration, and streaming are taught as mocked SDKs and brokers in Python. Docker, Kubernetes, Terraform, and a real AWS account are outside this tab, same as they are outside any browser-based course. See the roadmap page for exactly what runs here versus what is homework on your own machine.
Ready to start Day 1?
Core Python and SQL foundational modules are free to explore.