Learn · python
The language you will write almost every pipeline in.
Start from zero: what data engineering is, Python basics, file handling, APIs, and a production ingest capstone. No prior experience needed.
Core foundational modules are free. Advanced production modules need Pro.
A company has data in a hundred places: a CSV a vendor emails every morning, a payments API, a database another team owns. Somebody has to fetch it, fix the broken rows, and put it somewhere useful, on a schedule, without being woken up at 3 a.m. every time it fails. Doing that by hand does not scale past a week. Python is how you write it down once and let a machine repeat it.
Most data engineering job posts list Python first. In practice you use a small slice of it constantly: read a file, loop over records, call an API, handle the error, write a log line, exit with the right code. This track teaches that slice properly instead of teaching the whole language badly.
Nothing about programming
This track starts at what a variable is. If you can use a computer and type, you have enough.
Nothing about data engineering
The first lesson explains what the job actually is before any code appears.
5 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.
Before writing code, understand what a data engineer is paid to do and why Python is the tool. Skipping this is why beginners write scripts with no idea what problem they serve.
Getting StartedFree
0/2
What data engineering is, why Python is the language of choice, and how this track works.
By the end of this stage
You can explain in your own words what a pipeline is and what part of it Python handles.
Every transformation you will ever write is built from four things: values, decisions, repetition, and functions. Get these solid and the rest of the track is just vocabulary.
FoundationsFree
0/2
Variables, types, data structures, and the building blocks of every Python program.
Module Checkpoint
2 conceptual questions to verify mastery
Flow, functions & files
0/5
Loops, retries, reusable utilities, and reading warehouse files safely.
By the end of this stage
You can read a CSV off disk, loop through it, clean each row, and write the result back out.
Where people get stuck
Dictionaries and lists of dictionaries are the shape almost all pipeline data takes. If that shape still feels fuzzy, redo the exercises before moving on. Everything later assumes it.
Files on your laptop are the easy case. Real sources are HTTP APIs that paginate, time zones that lie, and fields that are sometimes null and sometimes the string 'null'.
APIs, streams & objectsFree preview
0/3
Talk to services, stream huge files, and model connectors as classes.
Module Checkpoint
1 conceptual questions to verify mastery
Time & validation
0/2
UTC intervals and rejecting dirty rows before they hit the lake.
By the end of this stage
You can pull data from a REST API page by page, parse timestamps correctly, and reject records that fail validation instead of silently corrupting the output.
Where people get stuck
Time zones and naive datetimes cause more production bugs than any other single thing in this stage. Slow down on that lesson.
A script that works once is not a pipeline. A pipeline runs unattended, fails safely, retries the transient errors, and does not double-write when it runs twice.
Text & PatternsFree preview
0/2
Parse logs and survive encoding surprises at the ingest boundary.
Reliable Pipelines
0/4
Logging, cleanup, retries, and an honest mental model of concurrency.
Module Checkpoint
1 conceptual questions to verify mastery
Packaging & Config
0/3
Twelve-factor config, pinning, and a CLI entrypoint, including what this browser cannot do.
By the end of this stage
You can write a job with logging, retries with backoff, config separated from code, and idempotent behaviour on rerun.
Where people get stuck
Idempotency is the concept most people nod along to and then get wrong. If you cannot explain what happens when your job runs twice on the same input, you have not got it yet.
The difference between a script and something a team will let near production is usually a test suite and a clean entrypoint.
Testing & Capstone
0/2
Prove transforms with tests, then assemble an ingest script that uses the whole track.
By the end of this stage
You finish with an ingest job you can put in a repo and talk about in an interview: it takes a real source, validates, transforms, loads, logs, and has tests.
Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.
30 minutes a day
about 10 sessions
1 hour a day
about 5 sessions
4 hours a weekend day
about 2 sessions
Do not batch this track into long sessions. One lesson plus its exercise per sitting beats four lessons read passively. If you are new to programming, expect the Foundations stage to take twice as long as the estimate, and that is normal.
These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.
| Common mistake | What to do instead |
|---|---|
| Learning Python from a general tutorial that spends three weeks on classes and decorators. | Data engineering Python is mostly files, dicts, requests, and error handling. Learn that first, and go deeper only when a real problem needs it. |
| Reading lessons without running the exercises. | The exercises are where the learning happens. Reading code gives you recognition, writing it gives you recall, and interviews test recall. |
| Catching every exception with a bare except and moving on. | That turns a loud failure into silent data loss. Catch the error you expect, and let the ones you did not expect crash the job so you find out. |