Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Core Python for Data Engineers

Progress0/42
x

Getting Started

  • What is data engineering?10m
  • Why Python for data engineers?8m

Foundations

  • Variables, types & type hints8m
  • DE data structures12m
  • Decisions and loops: if, for and while12m

Flow, functions & files

  • Control flow & error handling10m
  • Functions, modules & imports10m
  • Strings and text10m
  • File I/O & data formats12m
  • Working with JSON10m

APIs, streams & objectsPreview

  • Working with REST APIs12m
  • Iterators & generators (yield)Free12m
  • OOP for pipeline engineering12m

Time & validation

  • Working with dates & timestamps10m
  • Data validation with Pydantic12m

Text & PatternsPreview

  • Regular expressions for logsFree12m
  • String encoding & unicode gotchas12m

Reliable Pipelines

  • Logging instead of print-debugging12m
  • Context managers & resource cleanup10m
  • Retries, backoff, and idempotency14m
  • Concurrency, asyncio, and the GIL14m

Packaging & Config

  • Config & secrets management10m
  • Dependency management & pinning10m
  • Building a pipeline CLI12m

Testing & Capstone

  • Unit testing data transforms12m
  • Capstone: ingest script end to end18m

Databases, Cloud & Profiling

  • Databases from Python: sqlite3 & SQL engines15m
  • Cloud SDK from Python: S3 and object stores15m
  • Profiling: timeit, cProfile & memory tracking15m

DSA for Data Engineers

  • Big-O complexity & measurement15m
  • Hash maps & frequency counting15m
  • Two pointers & in-place array scanning15m
  • Sliding window & stream buffers15m
  • Prefix sums & range aggregates15m
  • Sorting & binary search with bisect15m
  • Stacks, queues & monotonic stacks15m
  • Heaps & priority queues with heapq15m
  • Linked lists & pointer chains15m
  • Trees, BST & hierarchical traversals15m
  • BFS, DFS & DAG traversals15m
  • Union-find & connected components15m
  • Dynamic programming fundamentals15m
Back to track
  1. Learn
  2. Core Python for Data Engineers
  3. Foundations
  4. Decisions and loops: if, for and while

Lesson 5 of 42 · Theory first, then run it

Decisions and loops: if, for and while

pythonbeginner12 min

Contents

  1. What you will do
  2. Why this skill
  3. How the code works
  4. Worked examples
  5. Common beginner questions
  6. What comes next
  7. Practice

Overview

if chooses a path, for goes through a list, while repeats until a condition is false. The pattern: start a total at zero, loop over rows, add when a condition is true.

On this page7 sections›
  1. 01What you will do
  2. 02Why this skill
  3. 03How the code works
  4. 04Worked examples
  5. 05Common beginner questions
  6. 06What comes next
  7. 07Practice

What you will do

Open a real orders file and you will not find 500 identical rows.

Some orders are paid, some are cancelled, and a few contain missing or negative amounts.

What happens when an orders file has different statuses?

When you download a daily export from a store, rows arrive in mixed states.

If your pipeline treats every row the same way, it either crashes or quietly corrupts financial reports.

  1. Incoming orderspaid, cancelled, corrupt
  2. ↓
  3. →
  4. Condition check (if)routes each row
  5. ↓
  6. →
  7. Clean reportsaccurate totals

Conditionals let code choose what to do with a record, while loops repeat the decision across every row.

An if statement inspects one row, while a for loop applies that rule to thousands of rows automatically.

Write the decision rule once. Let Python repeat it across every row.

Which loop helpers will you use?

Data pipelines require specific helpers to number rows, unpack dictionaries, and handle retries cleanly.

Rather than writing manual counter variables, Python standard helpers give you clean tools for each task.

The core looping tools for data pipelines.

HelperWhat it providesCommon pipeline use
enumerate()Row position and record togetherLogging exact line numbers of failed rows
.items()Dictionary keys and values togetherIterating config settings or summary maps
range()Sequence of fixed repeat countsBatch chunking and loop indexing
whileRuns until a condition turns falsePolling an API until status is completed
break & continueHalt loop or skip to next turnShort-circuit on match or bypass malformed records

Each helper eliminates manual counter variables and keeps pipeline logic concise and readable.

Mastering these helpers allows you to write robust parsers that handle edge cases gracefully.

Why this skill

Pipelines process millions of records that arrive on varying schedules.

Why can we not add up 2 lakh orders by hand or with copy-paste?

Leadership asks for yesterday paid revenue and cancelled count from 2 lakh orders.

Writing 2 lakh hardcoded lines breaks the moment tomorrow file arrives with a different row count.

Manual row handling vs loops
Manual / HardcodedLoop with conditionsscale200,000 lines of repeated codeBreaks when tomorrow has 250,000 rowsHigh chance of copy-paste mistakesOne compact loop handles any row countAdapts automatically when file sizechangesConsistent business logic across allrows

Pipelines process millions of records with constant code size.

A compact loop applies the exact same business logic whether a file has 10 rows or 10 million rows.

When tomorrow batch arrives with extra records, the loop processes them without code modifications.

In addition, iterating record by record consumes minimal computer memory compared to loading uncompressed spreadsheets.

A loop decouples pipeline logic from the number of incoming rows.

How the code works

Python gives you clear control structures to evaluate conditions and iterate collections.

How does Python make decisions with if, elif, and else?

An if statement tests a condition, executing indented lines only when the condition evaluates to True.

Add elif for secondary checks and else for remaining cases.

PythonBranching on order status
status = "paid"

  if status == "paid":
      print("process settlement")
  elif status == "refunded":
      print("reverse payment")
  else:
      print("flag for manual review")

Predict which branch prints when status is paid before checking the output below.

process settlement

The first condition was true, so Python executed that branch and skipped all remaining checks.

Using elif ensures only one branch runs, whereas separate if statements would test every single check independently.

What is the difference between == and =?

Two equals signs ask whether two values match. A single equals sign assigns a value.

Using a single equals sign in an if condition triggers an immediate syntax error.

The error for a single equals sign
SyntaxError: invalid syntax. Maybe you meant '==' or ':=' instead of '='?

Python warns you that an assignment was attempted where an equality comparison was expected.

Essential comparison and logical operators.

OperatorWhat it asksPipeline example
==Are both sides equal?status == 'paid'
!=Are they different?status != 'cancelled'
> <Bigger or smaller?total > 1000
>= <=Bigger or equal, smaller or equal?retries <= 3
andAre both conditions true?total > 0 and status == 'paid'
orIs at least one condition true?city == 'Mumbai' or city == 'Pune'
notInvert boolean valuenot is_duplicate

Joining conditions with and and or allows you to build sophisticated data validation filters.

Remember that comparisons evaluate from left to right, and parentheses clarify complex logic.

What happens if you forget indentation after a colon?

Python requires indented lines under every colon to know where branches start and stop.

Forgetting the four-space indentation halts execution immediately with an error.

The error for a missing indent
IndentationError: expected an indented block after 'if' statement on line 1

Notice that Python points directly to the line where an indented block was expected.

Every branch line ends with a colon and must be indented by four spaces.

How does a for loop read a list of records?

Most pipeline datasets arrive as lists of dictionaries, where each dictionary represents one record.

The for loop pulls one record at a time, making it accessible through a named variable.

PythonOne turn of the loop per order
orders = [
      {"customer": "Ravi", "total": 500},
      {"customer": "Anu", "total": 750},
      {"customer": "John", "total": 300},
  ]

  for order in orders:
      print(order["customer"], order["total"])
Ravi 500
  Anu 750
  John 300

On each turn, the order variable binds to the next dictionary until all records have been processed.

You extract specific fields using dictionary keys, such as order['total'].

How do we number rows using enumerate?

In data engineering, you frequently need both row index and record content to report error line numbers.

The enumerate helper supplies both, and start=1 sets human-readable counting.

PythonNumbering orders starting at 1
for position, order in enumerate(orders, start=1):
      print(position, order["customer"])
1 Ravi
  2 Anu
  3 John

Predict what prints if you omit start=1: does the first row number 0 or 1?

Default 0-based numbering without start=1
0 Ravi
  1 Anu
  2 John

Without start=1, Python starts numbering from index 0.

When a malformed record appears on row 482, enumerate lets your log statement pinpoint the exact line.

How do we inspect dictionary keys and values with .items()?

Iterating a dictionary directly yields only keys. Adding .items() yields key and value together.

This is standard practice when reading environment variables or summary metric tables.

PythonLooping with .items()
targets = {"Bengaluru": 100, "Mumbai": 150}

  for city, target in targets.items():
      print(city, target)
Bengaluru 100
  Mumbai 150

Two variables unpack the pair on every turn, perfect for reading configuration files and metric summaries.

If you only need values, Python provides .values(), but .items() delivers the complete picture.

How does range() repeat an action a fixed number of times?

When you need a loop to execute N times without an existing list, use range.

This is common when splitting large datasets into fixed batch chunks.

PythonRepeating 3 times
for i in range(3):
      print("Hello", i)

Predict the printed numbers before looking below: does it start at 0 or 1?

Hello 0
  Hello 1
  Hello 2

range(3) produces three numbers starting at 0 and stopping before 3.

To produce numbers from 1 to 3, pass two arguments: range(1, 4).

How do we write a bounded retry loop with while?

A for loop iterates known collections, while a while loop repeats as long as a condition remains True.

Always include an internal counter advancement so retry loops terminate safely.

PythonA retry loop with a limit
attempt = 0

  while attempt < 3:
      attempt = attempt + 1
      print("try number", attempt)

  print("stopped after", attempt, "tries")
try number 1
  try number 2
  try number 3
  stopped after 3 tries

When polling external APIs, bounded while loops prevent stalled jobs from hanging indefinitely.

Forgetting to advance the counter leaves the condition True forever, creating a runaway infinite loop.

Every while loop must advance its own termination condition.

How do continue and break control iteration?

Use continue to skip remaining steps for a bad row and break to exit the loop upon finding a match.

Here is a batch where one order has a missing total. We want the id of the first order exceeding 500.

PythonSkip bad rows, stop at the first match
orders = [
      {"id": 101, "customer": "Aarav", "total": 450.0},
      {"id": 102, "customer": "Priya", "total": None},
      {"id": 103, "customer": "Rahul", "total": 850.0},
      {"id": 104, "customer": "Neha", "total": 120.0},
  ]

  first_high = None
  for order in orders:
      if order["total"] is None:
          continue
      if order["total"] > 500:
          first_high = order["id"]
          break

  print("first high value order:", first_high)

Predict which order id prints, and whether order 104 is ever evaluated.

first high value order: 103

Order 102 was skipped by continue, order 103 triggered break, and order 104 was never evaluated.

Without the continue check, comparing None with a number triggers an immediate crash:

What happens without the continue check
TypeError: '>' not supported between instances of 'NoneType' and 'int'

Checking for None before comparisons protects pipelines against malformed inputs.

In production, continue filters out bad rows while break short-circuits search operations.

Worked examples

Now combine conditionals and loops into standard data engineering patterns.

How do we compute multiple metrics in a single pass?

Initialize totals at zero before the loop, then route rows into their respective accumulators.

Starting totals at 0.0 establishes floating-point arithmetic from the start.

PythonTwo totals in one pass
orders = [
      {"id": 101, "customer": "Aarav", "status": "paid", "total": 500.0},
      {"id": 102, "customer": "Priya", "status": "cancelled", "total": 300.0},
      {"id": 103, "customer": "Rahul", "status": "paid", "total": 120.5},
  ]

  paid_total = 0.0
  cancelled = 0

  for order in orders:
      if order["status"] == "paid":
          paid_total = paid_total + order["total"]
      elif order["status"] == "cancelled":
          cancelled = cancelled + 1

  print("paid total:", paid_total)
  print("cancelled:", cancelled)

Predict the final paid revenue and cancelled count before checking the output below.

paid total: 620.5
  cancelled: 1

paid_total and cancelled after each order.

After turnOrder looked atBranch that ranpaid_totalcancelled
Before loopnone yetnone0.00
1101 Aarav, paid, 500.0if (paid)500.00
2102 Priya, cancelled, 300.0elif (cancelled)500.01
3103 Rahul, paid, 120.5if (paid)620.51

Priya order was cancelled, so her 300 amount never touched paid_total; only the cancelled count incremented.

Tracing variable states turn by turn reveals exactly how conditions partition metrics.

How do we quarantine invalid rows instead of crashing?

When an order contains a negative or missing value, separate clean rows from bad rows using two lists.

This two-list approach prevents pipeline failure while preserving corrupted records for engineering review.

PythonSplitting clean rows from bad rows
orders = [
      {"id": 101, "customer": "Aarav", "total": 500.0},
      {"id": 102, "customer": "Priya", "total": -20.0},
      {"id": 103, "customer": "Rahul", "total": 120.5},
  ]

  good = []
  bad = []

  for order in orders:
      if order["total"] > 0:
          good.append(order)
      else:
          bad.append(order)

  print(len(good), "clean,", len(bad), "quarantined")
  print(bad)

Predict how many rows land in bad, and which customer order is quarantined.

2 clean, 1 quarantined
  [{'id': 102, 'customer': 'Priya', 'total': -20.0}]

Order 102 had a negative price, so it moved into quarantine whole without halting pipeline execution.

Never call orders.remove(order) during a loop; deleting elements while iterating shifts indices and skips records.

Building separate clean and quarantine lists operates in linear time without index corruption bugs.

How does a pipeline process rows from start to finish?

The complete loop and accumulator workflow
  1. 1

    Initialize accumulators: set totals and counters to zero outside the loop.

    1. paid_total = 0.0
    2. ↓
    3. →
    4. cancelled = 0
  2. 2

    Iterate records: inspect one dictionary per turn.

    1. for order in orders:
  3. 3

    Route with conditions: update matching totals and quarantine invalid rows.

    1. if paid
    2. ↓
    3. →
    4. elif cancelled
    5. ↓
    6. →
    7. else quarantine
  4. 4

    Deliver results: export clean totals and log quarantined exceptions.

    1. Reports
    2. ↓
    3. →
    4. Alert queue

Four rules to remember for loops in data engineering:

  1. Use for loops to iterate across collections and while loops for bounded polling retries.
  2. Always check for None before comparing values to prevent runtime TypeErrors.
  3. Accumulate metrics into initialized variables before the loop rather than recomputing them.
  4. Split malformed records into a quarantine list rather than modifying the iterated list in place.

Common beginner questions

When do I use for and when do I use while?

Use for when you have a list, dictionary, or fixed count. Use while when waiting for a condition to change, such as polling a database or job API.

Why does range(5) stop at 4?

Because it starts at 0 and stops before 5, giving exactly five numbers: 0, 1, 2, 3, 4. This matches 0-indexed list positions perfectly.

What is the difference between break and continue?

continue skips the rest of the current iteration and advances to the next item. break exits the entire loop immediately.

What comes next

In the next lesson you will learn how functions encapsulate looping logic into reusable, testable components.

You will also learn how try and except catch runtime errors so unexpected values never crash production pipelines.

Practice

In the exercise you get a list of customer transactions. Write a for loop that checks each status with if and elif.

Add up the revenue from paid orders, count cancelled orders, count unknown statuses, and store all three in a dictionary named result.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
DE data structuresControl flow & error handling