Overview
if chooses a path, for goes through a list, while repeats until a condition is false. The pattern: start a total at zero, loop over rows, add when a condition is true.
On this page7 sections
What you will do
Open a real orders file and you will not find 500 identical rows.
Some orders are paid, some are cancelled, and a few contain missing or negative amounts.
What happens when an orders file has different statuses?
When you download a daily export from a store, rows arrive in mixed states.
If your pipeline treats every row the same way, it either crashes or quietly corrupts financial reports.
- Incoming orderspaid, cancelled, corrupt
- Condition check (if)routes each row
- Clean reportsaccurate totals
Conditionals let code choose what to do with a record, while loops repeat the decision across every row.
An if statement inspects one row, while a for loop applies that rule to thousands of rows automatically.
Write the decision rule once. Let Python repeat it across every row.
Which loop helpers will you use?
Data pipelines require specific helpers to number rows, unpack dictionaries, and handle retries cleanly.
Rather than writing manual counter variables, Python standard helpers give you clean tools for each task.
The core looping tools for data pipelines.
| Helper | What it provides | Common pipeline use |
|---|---|---|
| enumerate() | Row position and record together | Logging exact line numbers of failed rows |
| .items() | Dictionary keys and values together | Iterating config settings or summary maps |
| range() | Sequence of fixed repeat counts | Batch chunking and loop indexing |
| while | Runs until a condition turns false | Polling an API until status is completed |
| break & continue | Halt loop or skip to next turn | Short-circuit on match or bypass malformed records |
Each helper eliminates manual counter variables and keeps pipeline logic concise and readable.
Mastering these helpers allows you to write robust parsers that handle edge cases gracefully.
Why this skill
Pipelines process millions of records that arrive on varying schedules.
Why can we not add up 2 lakh orders by hand or with copy-paste?
Leadership asks for yesterday paid revenue and cancelled count from 2 lakh orders.
Writing 2 lakh hardcoded lines breaks the moment tomorrow file arrives with a different row count.
Pipelines process millions of records with constant code size.
A compact loop applies the exact same business logic whether a file has 10 rows or 10 million rows.
When tomorrow batch arrives with extra records, the loop processes them without code modifications.
In addition, iterating record by record consumes minimal computer memory compared to loading uncompressed spreadsheets.
A loop decouples pipeline logic from the number of incoming rows.
How the code works
Python gives you clear control structures to evaluate conditions and iterate collections.
How does Python make decisions with if, elif, and else?
An if statement tests a condition, executing indented lines only when the condition evaluates to True.
Add elif for secondary checks and else for remaining cases.
status = "paid"
if status == "paid":
print("process settlement")
elif status == "refunded":
print("reverse payment")
else:
print("flag for manual review")Predict which branch prints when status is paid before checking the output below.
process settlement
The first condition was true, so Python executed that branch and skipped all remaining checks.
Using elif ensures only one branch runs, whereas separate if statements would test every single check independently.
What is the difference between == and =?
Two equals signs ask whether two values match. A single equals sign assigns a value.
Using a single equals sign in an if condition triggers an immediate syntax error.
SyntaxError: invalid syntax. Maybe you meant '==' or ':=' instead of '='?
Python warns you that an assignment was attempted where an equality comparison was expected.
Essential comparison and logical operators.
| Operator | What it asks | Pipeline example |
|---|---|---|
| == | Are both sides equal? | status == 'paid' |
| != | Are they different? | status != 'cancelled' |
| > < | Bigger or smaller? | total > 1000 |
| >= <= | Bigger or equal, smaller or equal? | retries <= 3 |
| and | Are both conditions true? | total > 0 and status == 'paid' |
| or | Is at least one condition true? | city == 'Mumbai' or city == 'Pune' |
| not | Invert boolean value | not is_duplicate |
Joining conditions with and and or allows you to build sophisticated data validation filters.
Remember that comparisons evaluate from left to right, and parentheses clarify complex logic.
What happens if you forget indentation after a colon?
Python requires indented lines under every colon to know where branches start and stop.
Forgetting the four-space indentation halts execution immediately with an error.
IndentationError: expected an indented block after 'if' statement on line 1
Notice that Python points directly to the line where an indented block was expected.
Every branch line ends with a colon and must be indented by four spaces.
How does a for loop read a list of records?
Most pipeline datasets arrive as lists of dictionaries, where each dictionary represents one record.
The for loop pulls one record at a time, making it accessible through a named variable.
orders = [
{"customer": "Ravi", "total": 500},
{"customer": "Anu", "total": 750},
{"customer": "John", "total": 300},
]
for order in orders:
print(order["customer"], order["total"])Ravi 500 Anu 750 John 300
On each turn, the order variable binds to the next dictionary until all records have been processed.
You extract specific fields using dictionary keys, such as order['total'].
How do we number rows using enumerate?
In data engineering, you frequently need both row index and record content to report error line numbers.
The enumerate helper supplies both, and start=1 sets human-readable counting.
for position, order in enumerate(orders, start=1):
print(position, order["customer"])1 Ravi 2 Anu 3 John
Predict what prints if you omit start=1: does the first row number 0 or 1?
0 Ravi 1 Anu 2 John
Without start=1, Python starts numbering from index 0.
When a malformed record appears on row 482, enumerate lets your log statement pinpoint the exact line.
How do we inspect dictionary keys and values with .items()?
Iterating a dictionary directly yields only keys. Adding .items() yields key and value together.
This is standard practice when reading environment variables or summary metric tables.
targets = {"Bengaluru": 100, "Mumbai": 150}
for city, target in targets.items():
print(city, target)Bengaluru 100 Mumbai 150
Two variables unpack the pair on every turn, perfect for reading configuration files and metric summaries.
If you only need values, Python provides .values(), but .items() delivers the complete picture.
How does range() repeat an action a fixed number of times?
When you need a loop to execute N times without an existing list, use range.
This is common when splitting large datasets into fixed batch chunks.
for i in range(3):
print("Hello", i)Predict the printed numbers before looking below: does it start at 0 or 1?
Hello 0 Hello 1 Hello 2
range(3) produces three numbers starting at 0 and stopping before 3.
To produce numbers from 1 to 3, pass two arguments: range(1, 4).
How do we write a bounded retry loop with while?
A for loop iterates known collections, while a while loop repeats as long as a condition remains True.
Always include an internal counter advancement so retry loops terminate safely.
attempt = 0
while attempt < 3:
attempt = attempt + 1
print("try number", attempt)
print("stopped after", attempt, "tries")try number 1 try number 2 try number 3 stopped after 3 tries
When polling external APIs, bounded while loops prevent stalled jobs from hanging indefinitely.
Forgetting to advance the counter leaves the condition True forever, creating a runaway infinite loop.
Every while loop must advance its own termination condition.
How do continue and break control iteration?
Use continue to skip remaining steps for a bad row and break to exit the loop upon finding a match.
Here is a batch where one order has a missing total. We want the id of the first order exceeding 500.
orders = [
{"id": 101, "customer": "Aarav", "total": 450.0},
{"id": 102, "customer": "Priya", "total": None},
{"id": 103, "customer": "Rahul", "total": 850.0},
{"id": 104, "customer": "Neha", "total": 120.0},
]
first_high = None
for order in orders:
if order["total"] is None:
continue
if order["total"] > 500:
first_high = order["id"]
break
print("first high value order:", first_high)Predict which order id prints, and whether order 104 is ever evaluated.
first high value order: 103
Order 102 was skipped by continue, order 103 triggered break, and order 104 was never evaluated.
Without the continue check, comparing None with a number triggers an immediate crash:
TypeError: '>' not supported between instances of 'NoneType' and 'int'
Checking for None before comparisons protects pipelines against malformed inputs.
In production, continue filters out bad rows while break short-circuits search operations.
Worked examples
Now combine conditionals and loops into standard data engineering patterns.
How do we compute multiple metrics in a single pass?
Initialize totals at zero before the loop, then route rows into their respective accumulators.
Starting totals at 0.0 establishes floating-point arithmetic from the start.
orders = [
{"id": 101, "customer": "Aarav", "status": "paid", "total": 500.0},
{"id": 102, "customer": "Priya", "status": "cancelled", "total": 300.0},
{"id": 103, "customer": "Rahul", "status": "paid", "total": 120.5},
]
paid_total = 0.0
cancelled = 0
for order in orders:
if order["status"] == "paid":
paid_total = paid_total + order["total"]
elif order["status"] == "cancelled":
cancelled = cancelled + 1
print("paid total:", paid_total)
print("cancelled:", cancelled)Predict the final paid revenue and cancelled count before checking the output below.
paid total: 620.5 cancelled: 1
paid_total and cancelled after each order.
| After turn | Order looked at | Branch that ran | paid_total | cancelled |
|---|---|---|---|---|
| Before loop | none yet | none | 0.0 | 0 |
| 1 | 101 Aarav, paid, 500.0 | if (paid) | 500.0 | 0 |
| 2 | 102 Priya, cancelled, 300.0 | elif (cancelled) | 500.0 | 1 |
| 3 | 103 Rahul, paid, 120.5 | if (paid) | 620.5 | 1 |
Priya order was cancelled, so her 300 amount never touched paid_total; only the cancelled count incremented.
Tracing variable states turn by turn reveals exactly how conditions partition metrics.
How do we quarantine invalid rows instead of crashing?
When an order contains a negative or missing value, separate clean rows from bad rows using two lists.
This two-list approach prevents pipeline failure while preserving corrupted records for engineering review.
orders = [
{"id": 101, "customer": "Aarav", "total": 500.0},
{"id": 102, "customer": "Priya", "total": -20.0},
{"id": 103, "customer": "Rahul", "total": 120.5},
]
good = []
bad = []
for order in orders:
if order["total"] > 0:
good.append(order)
else:
bad.append(order)
print(len(good), "clean,", len(bad), "quarantined")
print(bad)Predict how many rows land in bad, and which customer order is quarantined.
2 clean, 1 quarantined
[{'id': 102, 'customer': 'Priya', 'total': -20.0}]
Order 102 had a negative price, so it moved into quarantine whole without halting pipeline execution.
Never call orders.remove(order) during a loop; deleting elements while iterating shifts indices and skips records.
Building separate clean and quarantine lists operates in linear time without index corruption bugs.
How does a pipeline process rows from start to finish?
Initialize accumulators: set totals and counters to zero outside the loop.
- paid_total = 0.0
- cancelled = 0
Iterate records: inspect one dictionary per turn.
- for order in orders:
Route with conditions: update matching totals and quarantine invalid rows.
- if paid
- elif cancelled
- else quarantine
Deliver results: export clean totals and log quarantined exceptions.
- Reports
- Alert queue
Four rules to remember for loops in data engineering:
- Use for loops to iterate across collections and while loops for bounded polling retries.
- Always check for None before comparing values to prevent runtime TypeErrors.
- Accumulate metrics into initialized variables before the loop rather than recomputing them.
- Split malformed records into a quarantine list rather than modifying the iterated list in place.
Common beginner questions
When do I use for and when do I use while?
Use for when you have a list, dictionary, or fixed count. Use while when waiting for a condition to change, such as polling a database or job API.
Why does range(5) stop at 4?
Because it starts at 0 and stops before 5, giving exactly five numbers: 0, 1, 2, 3, 4. This matches 0-indexed list positions perfectly.
What is the difference between break and continue?
continue skips the rest of the current iteration and advances to the next item. break exits the entire loop immediately.
What comes next
In the next lesson you will learn how functions encapsulate looping logic into reusable, testable components.
You will also learn how try and except catch runtime errors so unexpected values never crash production pipelines.
Practice
In the exercise you get a list of customer transactions. Write a for loop that checks each status with if and elif.
Add up the revenue from paid orders, count cancelled orders, count unknown statuses, and store all three in a dictionary named result.
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.