Overview
Lists, tuples, dicts, sets, and comprehensions for pipeline data.
On this page8 sections
How to choose a structure by asking a question
Do not start by asking which collection is "best." Start by asking what operation your pipeline needs to perform repeatedly. If the question is "process every row in arrival order," a list is natural. If the question is "find the row for this ID," a dictionary is natural. If the question is "have I seen this value before?" a set is often natural. If the question is "represent a fixed combination of values," a tuple may fit.
The performance difference is not just a computer-science detail. Data pipelines repeatedly perform the same operations over thousands or millions of records. A structure that matches the operation can remove entire loops from your program. That improves both runtime and readability because the code expresses the question directly.
Also separate two ideas that beginners often mix up: changing a container and changing the objects inside it. A new list can still contain references to the same dictionaries as the old list. This is why copying a container is not always the same as copying all nested data. You will meet this distinction again when pipeline stages share records.
What you will do
Python provides four built-in ways to organize collections of data: lists, tuples, dictionaries, and sets. Each one is designed for a different job. Choosing the right container makes your code faster, clearer, and less error-prone.
A list is an ordered collection that can grow, shrink, and change. A tuple is like a list that is locked after creation (immutable). A dictionary (dict) maps unique keys to values, like a phone book mapping names to numbers. A set is an unordered collection that automatically removes duplicates.
A list comprehension is a concise way to create a new list by transforming or filtering an existing collection. The syntax [expression for item in collection if condition] replaces several lines of loop code with one readable line.
Read a comprehension left to right once you know the pieces. The for item in collection part is the same loop you already write. The expression before for is what each item becomes. The optional if at the end is the filter. Start with a plain for loop if a comprehension feels opaque; both produce the same list, and clarity beats cleverness.
Why this skill
Imagine a payment processor sends your company 50,000 order records every hour. You need to look up individual orders by their ID, detect duplicate submissions, filter for paid orders, and compute totals. If you store all orders in a plain list and search it every time, each lookup scans up to 50,000 items. With a dictionary keyed by order_id, each lookup is nearly instant regardless of how many orders exist. Choosing the wrong data structure can turn a 1-second job into a 10-minute job.
How the code works
Think of a list like a numbered row of lockers: items have positions (index 0, 1, 2...) and you can add or remove lockers. A dictionary is like a filing cabinet with labeled folders: you find a folder instantly by its label without checking every drawer.
Lists preserve order. Dictionaries enable instant lookup by key. Sets eliminate duplicates. Tuples freeze data.
Pick the structure based on what question you need to answer about the data.
| Structure | Ordered? | Changeable? | Best use case | Lookup speed |
|---|---|---|---|---|
| list | Yes | Yes | Batch of rows to process in order | By index: instant. By value: slow (scans all) |
| tuple | Yes | No | Fixed record or compound key | By index: instant |
| dict | Insertion order | Yes | Find a row by its unique ID | By key: instant (average) |
| set | No | Yes | Remove duplicates or check membership | Membership check: instant (average) |
Here is sample order data we will organize into different structures:
Three orders with different statuses.
| id | status | total |
|---|---|---|
| ORD-1 | paid | 120.50 |
| ORD-2 | cancelled | 40.00 |
| ORD-3 | paid | 88.25 |
Worked examples
Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.
orders = [
{"id": "ORD-1", "status": "paid", "total": 120.5},
{"id": "ORD-2", "status": "cancelled", "total": 40.0},
{"id": "ORD-3", "status": "paid", "total": 88.25},
]
# List comprehension: filter paid totals
paid_totals = [row["total"] for row in orders if row["status"] == "paid"]
# Dictionary comprehension: index by order ID
by_id = {row["id"]: row for row in orders}
# Set comprehension: unique statuses
statuses = {row["status"] for row in orders}
print("Paid totals:", paid_totals)
print("Lookup ORD-3:", by_id["ORD-3"])
print("Unique statuses:", sorted(statuses))The list comprehension collects only the totals from paid orders. The dictionary comprehension builds an index so you can find any order by its ID instantly. The set comprehension collects all unique statuses, automatically removing duplicates.
One beginner trap is that a variable does not contain a separate copy of a list or dictionary; it points to an object. If two names point to the same list, changing it through either name changes the same batch. This matters when a transform is meant to leave its input untouched: make a copy or build a new list rather than accidentally changing data that another stage still needs.
raw_rows = [{"id": "ORD-1", "status": "paid"}]
# Both names point to the same list
working_rows = raw_rows
working_rows[0]["status"] = "loaded"
print(raw_rows[0]["status"]) # loaded: the original changed too
# Build a new row when you want an independent result
clean_rows = [{**row, "status": "clean"} for row in raw_rows]
print(clean_rows[0]["status"]) # cleanThe second example uses {**row, ...} to create a new dictionary for each row. That style is often safer in a pipeline: raw input remains available for auditing, while the next stage receives a clean result. In-place changes with row["status"] = ... are not wrong, but use them deliberately and document which stage owns the data.
Python uses a hash table internally. Given an order_id, it jumps directly to the matching row without scanning.
# Check for duplicate IDs before building a dictionary index
order_ids = [row["id"] for row in orders]
if len(order_ids) != len(set(order_ids)):
raise ValueError("Duplicate order_id found in batch!")
# Safe to build the index now
by_id = {row["id"]: row for row in orders}
# Use the index for instant lookups
print("Order ORD-1:", by_id["ORD-1"])
print("Order ORD-3:", by_id["ORD-3"])
# Compute paid total using filtered list
result = [row["total"] for row in orders if row["status"] == "paid"]
print("Paid totals:", result)
print("Sum:", sum(result))Dictionaries cannot have duplicate keys. If two orders share the same ID, the second silently overwrites the first. Always verify uniqueness before building a dictionary index from batch data.
# WRONG: scanning the entire list for each lookup (slow for large data)
for wanted_id in ["ORD-1", "ORD-3"]:
for order in orders:
if order["id"] == wanted_id:
print(order)
break
# RIGHT: build the index once, then look up instantly
by_id = {row["id"]: row for row in orders}
for wanted_id in ["ORD-1", "ORD-3"]:
print(by_id[wanted_id])The wrong approach scans the entire list for every lookup. With 50,000 orders and 1,000 lookups, that is 50 million comparisons. The right approach builds the dictionary once (50,000 operations) and then each lookup is instant. For large datasets, this difference is enormous.
Duplicate keys are silently overwritten
If you build a dictionary from data with duplicate keys, the last value wins with no error or warning. Always check len(ids) == len(set(ids)) before indexing.
KeyError means the key does not exist
Accessing by_id["ORD-999"] raises a KeyError if that key is missing. Use by_id.get("ORD-999") to get None instead of a crash, or check with 'if key in by_id' first.
A tuple's other job, beyond freezing a record, is standing in as a compound key: a dictionary key made of more than one value glued together.
orders = [
{"date": "2026-08-01", "product": "laptop", "total": 999.99},
{"date": "2026-08-01", "product": "mouse", "total": 24.99},
{"date": "2026-08-02", "product": "laptop", "total": 999.99},
]
# Key on (date, product) together, not just one field
by_day_and_product: dict[tuple[str, str], float] = {}
for row in orders:
key = (row["date"], row["product"])
by_day_and_product[key] = by_day_and_product.get(key, 0) + row["total"]
print(by_day_and_product[("2026-08-01", "laptop")])Revenue per product per day is a common report, and neither date alone nor product alone is a unique key for it. The pair (date, product) is. A list could not do this job even if you wanted it to: dictionary keys must be hashable, meaning Python can compute a fixed fingerprint for them, and only immutable values like tuples, strings, and numbers qualify. A list can change after creation, so its fingerprint could not stay valid, and Python refuses it as a key.
Under the hood: CPython hash tables and open addressing
CPython dictionaries and sets are implemented as compact open-addressed hash tables with a sparse index array and a dense entries array. When you query `key in dictionary`, Python hashes the key using `hash(key)`, maps it to an index modulo the table size, and checks the slot. In the absence of hash collisions, lookup is O(1). If two keys produce the same hash slot, Python probes alternate slots using a perturbation formula. Sets and dicts automatically resize when they reach two-thirds capacity (a load factor of 0.66).
Real data engineering usage: UPI reconciliation at scale
In daily banking reconciliation (such as matching PhonePe or Google Pay transaction logs against a bank core ledger), a pipeline receives 2 million records from the app gateway and 2 million records from the bank. If you use a list lookup (`for row in gateway: if row['tx_id'] in bank_list:`), searching a 2-million item list for each record triggers $2,000,000 \times 2,000,000 = 4 \times 10^{12}$ comparisons, freezing the pipeline for days. By indexing bank transaction IDs into a Python set (`bank_tx_set = {row['tx_id'] for row in bank_records}`), the membership check executes in O(1) time. The entire 2-million record reconciliation completes in under 2 seconds.
Common beginner confusion: Shallow copy versus deep copy traps
A classic junior engineer trap occurs when attempting to copy records: `working_batch = raw_records.copy()` or `list(raw_records)`. This creates a shallow copy. The outer list is fresh, but every dictionary inside the list points to the exact same memory address as the original. Modifying `working_batch[0]['status'] = 'processed'` silently mutates `raw_records[0]['status']` as well, destroying the raw source audit trail. When you need independent dictionaries, use dictionary comprehension copies (`[{**row} for row in raw_records]`) or `copy.deepcopy()`.
Interview connection: List versus set search complexity
An interviewer asks: 'What is the time complexity of searching for an item in a Python list versus a set, and why? Under what circumstances can a set lookup degrade?' A strong answer: List lookup is O(N) because it performs a linear sequential scan from index 0 to N. Set lookup is O(1) average time because it hashes the key and directly inspects the hash bucket. It can degrade to O(N) only in the pathological scenario where every key has the exact same hash collision, forcing Python to linearly probe every entry.
Choose the container for the access pattern
Use lists when order and iteration matter; use sets for uniqueness and instant membership checks; use dictionaries for instant key-based row lookups.
Common beginner questions
When should I use a list versus a dictionary?
Use a list when you need to process items in order, like iterating over all orders, or when position matters. Use a dictionary when you need to look up items by a unique identifier, like finding an order by its ID.
What is the difference between a list and a tuple?
A list can be changed after creation: you can add items, remove items, or modify items. A tuple cannot be changed once created. Use tuples for data that should never change, like a date represented as (2026, 8, 1), or a compound key like ("US", "CA").
Does the order I loop over a dictionary come back the same every time?
Yes, in modern Python a dictionary remembers the order keys were inserted, and looping over it returns keys in that order. It is still not the same as a list: you cannot get the third item by position without extra work, since dictionaries are built for lookup by key, not by index.
Why not just use a list for everything?
Lists work fine for small datasets. But as your data grows, the wrong structure causes real performance problems. Looking up one item in a list of 1 million rows means scanning up to 1 million items. A dictionary does it in one step.
What comes next
You now know how to organize data into the right containers. The next lesson covers control flow: making decisions with if/else, repeating work with loops, and handling errors gracefully.
Practice
Run Sample to see the dictionary lookup, unique statuses, and paid totals. Then complete the Exercise: build a dictionary keyed by id, collect paid totals into a list called result that equals [120.5, 88.25], and print it.
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.
Practice this
Same ideas as interview drills. These challenges open in the studio with a dataset and tests already set up.
- Contains DuplicateInterview-style drill: True if any value appears more than once.Studiobeginnerpython10 min
- Find IntersectionInterview-style drill: Return the sorted unique intersection of two lists.Studiobeginnerpython10 min
- First Occurrence MapInterview-style drill: Map each distinct character to the index of its first occurrence.Studiobeginnerpython10 min