Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Core Python for Data Engineers

Progress0/25
x

Getting Started

  • What is data engineering?10m
  • Why Python for data engineers?8m

Foundations

  • Variables, types & type hints8m
  • DE data structures12m

Flow, functions & files

  • Control flow & error handling10m
  • Functions, modules & imports10m
  • Strings and text10m
  • File I/O & data formats12m
  • Working with JSON10m

APIs, streams & objectsPreview

  • Working with REST APIs12m
  • Iterators & generators (yield)Free12m
  • OOP for pipeline engineering12m

Time & validation

  • Working with dates & timestamps10m
  • Data validation with Pydantic12m

Text & PatternsPreview

  • Regular expressions for logsFree12m
  • String encoding & unicode gotchas12m

Reliable Pipelines

  • Logging instead of print-debugging12m
  • Context managers & resource cleanup10m
  • Retries, backoff, and idempotency14m
  • Concurrency, asyncio, and the GIL14m

Packaging & Config

  • Config & secrets management10m
  • Dependency management & pinning10m
  • Building a pipeline CLI12m

Testing & Capstone

  • Unit testing data transforms12m
  • Capstone: ingest script end to end18m
Back to track
  1. Learn
  2. Core Python for Data Engineers
  3. APIs, streams & objects
  4. Iterators & generators (yield)

Lesson 11 of 25 · Theory first, then run it

Iterators & generators (yield)

pythonintermediate12 min

Overview

Process one line at a time with yield instead of loading the whole file.

On this page8 sections›
  1. 1The mental model: produce, pause, resume
  2. 2What you will do
  3. 3Why this skill
  4. 4How the code works
  5. 5Worked examples
  6. 6Common beginner questions
  7. 7What comes next
  8. 8Practice

The mental model: produce, pause, resume

The word "lazy" in Python means the work is delayed until a value is requested. A generator therefore represents a process that can produce a sequence over time. It does not mean the work disappears; it means the program does not perform all of that work upfront.

Imagine a function reaching yield. At that moment, Python gives the yielded value to the caller and preserves the function's execution state. When the caller asks for another value, execution resumes after the yield and continues until the next yield or until the function finishes. That pause-and-resume behavior is the heart of generators.

The biggest data-engineering benefit is memory control. If a pipeline only needs to inspect one row, transform it, and pass it onward, keeping millions of completed rows in a list is unnecessary. A generator lets the pipeline keep only the small amount of state required to produce the next value.

What you will do

A generator is a special kind of function that produces values one at a time instead of computing them all at once. Instead of using return to send back one final result, a generator uses yield to produce a value, pause, and wait until the caller asks for the next one.

When you call a regular function, it runs completely and returns. When you call a generator function, it returns a generator object immediately without running any code. Each time you request the next value (by calling next() or using a for loop), the generator runs until it hits yield, produces that value, and pauses until you ask again.

The key benefit: generators do not build the entire result in memory. They compute values on demand. This means you can process a 10 GB file while using only a few kilobytes of memory, because only one line exists in memory at a time.

Why this skill

Imagine a data pipeline that processes daily access logs. Each log file is 12 GB. If you read the entire file into a list with readlines(), you need 12 GB of RAM just to hold the raw text. On a server with 8 GB of RAM, the program crashes with an out-of-memory error. With a generator that yields one line at a time, memory usage stays constant regardless of file size. You can process a 12 GB file or a 120 GB file with the same few kilobytes of working memory.

How the code works

Think of a generator like a bookmark in a book. Instead of photocopying the entire book (loading it all into memory), you read one page at a time, place your bookmark, and come back later for the next page. The book stays on the shelf; only one page is in your hands at any moment.

One line at a time through the processor
Open fileRead one lineyield strippedlineCaller processesitRequest nextline

The generator reads a line, yields it for processing, then forgets it and moves to the next. Memory stays flat.

Choose a generator when you process data sequentially and do not need to go back.

PropertyList (eager)Generator (lazy)
When values are computedAll at once, immediatelyOne at a time, on demand
Memory usageProportional to data sizeConstant (just the current value)
Can access by index?Yes: items[5]No: must iterate in order
Can iterate twice?YesNo: exhausted after one pass
Best forSmall datasets you reuseLarge files or streams
Memory comparison: list versus generator
5Generator (1line inmemory)95List (alllines inmemory)

For a 40-line file the difference is small. For 40 million lines, the generator prevents crashes.

Worked examples

Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.

PythonCreating and advancing a generator
from collections.abc import Iterator

def lines(path: str) -> Iterator[str]:
    """Yield one stripped line at a time from a file."""
    with open(path, encoding="utf-8") as f:
        for line in f:
            yield line.rstrip()

# Create the generator (no lines read yet)
stream = lines("/data/access.log")

# Pull two lines manually
print("Line 1:", next(stream))
print("Line 2:", next(stream))
print("(remaining lines are still unread on disk)")

The lines function looks like a normal function but uses yield instead of return. When called, it returns a generator object without reading any lines. Each call to next(stream) runs the function until the next yield, produces that line, and pauses. The file is read incrementally, one line at a time.

PythonCounting 40 lines without materializing the file
def lines(path: str):
    """Yield stripped lines from a file."""
    with open(path, encoding="utf-8") as f:
        for line in f:
            yield line.rstrip()

# Count all lines without loading the file into memory
result = 0
for line in lines("/data/access.log"):
    result += 1
    if result <= 3:
        print(f"Sample line {result}: {line}")

print("Total line count:", result)

The for loop consumes the generator one value at a time. At no point does the entire file exist in memory as a list. The result variable counts lines as they pass through. For the first three lines, we also print them as a sample. The fixture file has 40 lines.

PythonDo not materialize a generator into a list
def lines(path: str):
    with open(path, encoding="utf-8") as f:
        for line in f:
            yield line.rstrip()

# WRONG: calling list() defeats the purpose of a generator
all_lines = list(lines("/data/access.log"))
print(len(all_lines))  # works but loads everything into memory

# RIGHT: consume the stream incrementally
count = 0
for _ in lines("/data/access.log"):
    count += 1
print(count)  # same result, constant memory

Calling list() on a generator forces it to produce every value at once and store them all in a list. This completely defeats the memory benefit. Only materialize (convert to a list) when you are certain the data fits in memory and you need random access or multiple passes.

The pattern that makes generators genuinely useful is chaining them. Each stage takes a stream in and yields a stream out, so a whole pipeline runs with one row in memory at a time. Here is a realistic version: an access log where a payments team needs the count and total of server errors.

Input: plain text, one request per line, with the status code as the fourth field.

one line of /data/access.log
2026-08-01T08:00:01Z GET /api/orders 200 0.031
2026-08-01T08:00:02Z GET /api/orders 503 1.204
2026-08-01T08:00:04Z POST /api/orders 500 0.887
PythonA business scenario: three chained generators over one log file
from collections.abc import Iterator

def lines(path: str) -> Iterator[str]:
    """Stage 1: yield one raw line at a time."""
    with open(path, encoding="utf-8") as f:
        for line in f:
            yield line.rstrip()

def parsed(rows: Iterator[str]) -> Iterator[dict]:
    """Stage 2: turn each line into a dict."""
    for line in rows:
        parts = line.split()
        if len(parts) < 4:
            continue                      # skip malformed lines
        yield {"ts": parts[0], "path": parts[2], "status": parts[3]}

def errors_only(rows: Iterator[dict]) -> Iterator[dict]:
    """Stage 3: keep only 5xx responses."""
    for row in rows:
        if row["status"].startswith("5"):
            yield row

# Nothing has been read yet. The pipeline is just wired up.
pipeline = errors_only(parsed(lines("/data/access.log")))

count = 0
for row in pipeline:
    count += 1

print("Server errors found:", count)

Trace one line through that chain and the laziness becomes concrete. The for loop asks the pipeline for a row. errors_only asks parsed for a row. parsed asks lines for a row. lines reads exactly one line off disk and yields it back up. If that line is not a 5xx, errors_only discards it and asks for another. At no point does a list of all lines, or all parsed dicts, exist anywhere. Memory stays flat whether the file has 40 lines or 40 million.

Two limits are worth being upfront about. First, this reads the file once and only forward, so you cannot ask "what was the previous row" without deliberately keeping it in a variable, and you cannot loop over pipeline a second time without rebuilding it. Second, laziness makes errors arrive late: a broken line in the middle of a 12 GB file raises its exception after ten minutes of processing, not at the moment you called the function. Both are the price of constant memory, and for large files it is usually worth paying.

Under the hood: Generator state machines and frame suspension

When a normal function returns, CPython pops its execution frame (`PyFrameObject`) off the call stack and destroys its local namespace. A generator function behaves differently: encountering `yield` pauses execution, writes the yielded object to the caller, and leaves the stack frame allocated on the Python heap. The bytecode instruction pointer (`f_lasti`) records the exact position. When `next()` is called, CPython resumes execution at `f_lasti + 1` with all local variables preserved. Because only one frame is active and no collection is accumulated, memory consumption remains strictly $O(1)$.

Real data engineering usage: Server-side database streaming

When extracting 50 million rows from an OLTP database (such as PostgreSQL or MySQL) into a Parquet lakehouse file, running `cursor.fetchall()` pulls all 50 million rows into application RAM, crashing the server. Instead, data engineers use server-side cursors via generators (`psycopg2.extras.DictCursor` with a declared cursor name and `cursor.itersize = 10000`). The database streams rows in 10,000-row chunks over the network wire, while the Python generator yields individual rows directly to a Parquet writer, processing terabytes of data within a 512 MB memory boundary.

Common beginner confusion: Attempting to re-iterate or index generators

A junior engineer once wrote: `orders = stream_orders(); print(len(orders)); for o in orders: process(o)`. This code fails twice: first, `len(orders)` raises `TypeError: object of type 'generator' has no len()`; second, if they had used `list(orders)` to find the length, the generator would be completely exhausted, leaving zero rows for the subsequent loop. A generator is a single-use stream. Once consumed to the end, it cannot be reset without calling the generator function again.

Interview connection: Generators vs list comprehensions in pipelines

An interviewer asks: 'What is the trade-off between a generator expression `(x for x in data)` and a list comprehension `[x for x in data]` in a data processing pipeline?' A strong response: A list comprehension is eager: it allocates memory immediately and stores all elements in RAM ($O(N)$ space), enabling fast repeated indexing and sorting. A generator expression is lazy: it produces elements on demand in $O(1)$ space, making it mandatory for massive files and streams. However, generators cannot be indexed, cannot report length without exhaustion, and can only be traversed once.

Generators produce streams, not storage

Use generators to build streaming pipelines that process one record at a time. Never materialize an entire stream into a list unless you have proven the dataset fits comfortably in RAM.

A generator is consumed after one pass

After a for loop finishes, the generator is exhausted and empty. You cannot iterate over it again. If you need a second pass, call the generator function again to create a fresh generator.

Generators still do O(n) work

A generator does not make processing faster. It makes memory usage constant. You still read and process every single line. The benefit is memory efficiency, not speed.

Common beginner questions

When should I use a generator instead of a list?

Use a generator when the data might not fit in memory, when one forward pass is all you need, and when you never index into it. Use a list when the data is small and you want to walk it more than once, sort it, or grab items[500]. If you are unsure and the data is small, a list is simpler and there is no shame in it.

What happens if I call next() after the generator is done?

Python raises StopIteration. A for loop catches that automatically, which is why loops just end quietly. If you call next() by hand, either wrap it in try/except StopIteration or pass a default: next(gen, None) hands back None instead of raising.

Can I use a generator with functions like sum() or max()?

Yes, and this is where generators read beautifully. sum(), max(), min(), any(), all(), and sorted() all accept one. sum(int(line) for line in lines(path)) totals a file without ever building an intermediate list. Note that sorted() is the exception that must hold everything in memory to do its job, because you cannot sort what you have not seen yet.

Does using a generator make my code faster?

Usually not, and it is worth setting that expectation. You still read and process every row, so the total work is the same. What changes is peak memory, which stays flat instead of growing with the file. The speed-up you sometimes see is indirect: a job that no longer swaps to disk or crashes and restarts finishes sooner.

What comes next

Generators help you handle large data efficiently. The next lesson introduces object-oriented programming (classes), which lets you bundle data and behavior into reusable, swappable components.

Practice

Run Sample and notice that next() returns only two lines without reading the rest of the file. Complete the Exercise: write a lines(path) generator function using yield, count all lines in /data/access.log using a for loop (without list()), assign 40 to result, and print it.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Practice this

Same ideas as interview drills. These challenges open in the studio with a dataset and tests already set up.

  • Best 3-day sales windowInterview-style drill: Given daily sales figures, find the 3-day consecutive window with the highest total.Studiointermediatepython15 minPro
  • Find All AnagramsInterview-style drill: Start indices of every anagram of p inside s.Studiointermediatepython14 minPro
  • Find Anagrams WindowInterview-style drill: Same behavior as find_anagrams: start indices of anagram windows.Studiointermediatepython14 minPro
Rate:
Was this useful?
Working with REST APIsOOP for pipeline engineering