Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Core Python for Data Engineers

Progress0/25
x

Getting Started

  • What is data engineering?10m
  • Why Python for data engineers?8m

Foundations

  • Variables, types & type hints8m
  • DE data structures12m

Flow, functions & files

  • Control flow & error handling10m
  • Functions, modules & imports10m
  • Strings and text10m
  • File I/O & data formats12m
  • Working with JSON10m

APIs, streams & objectsPreview

  • Working with REST APIs12m
  • Iterators & generators (yield)Free12m
  • OOP for pipeline engineering12m

Time & validation

  • Working with dates & timestamps10m
  • Data validation with Pydantic12m

Text & PatternsPreview

  • Regular expressions for logsFree12m
  • String encoding & unicode gotchas12m

Reliable Pipelines

  • Logging instead of print-debugging12m
  • Context managers & resource cleanup10m
  • Retries, backoff, and idempotency14m
  • Concurrency, asyncio, and the GIL14m

Packaging & Config

  • Config & secrets management10m
  • Dependency management & pinning10m
  • Building a pipeline CLI12m

Testing & Capstone

  • Unit testing data transforms12m
  • Capstone: ingest script end to end18m
Back to track
  1. Learn
  2. Core Python for Data Engineers
  3. Foundations
  4. Variables, types & type hints

Lesson 3 of 25 · Theory first, then run it

Variables, types & type hints

pythonbeginner8 min

Overview

Name a value, pick its type, and cast text before you do math.

On this page8 sections›
  1. 1The mental model: names, objects, and types
  2. 2What you will do
  3. 3Why this skill
  4. 4How the code works
  5. 5Worked examples
  6. 6Common beginner questions
  7. 7What comes next
  8. 8Practice

The mental model: names, objects, and types

A beginner-friendly way to start is to imagine a variable as a label attached to a value. But remember that this is a learning model, not the complete implementation detail. In Python, names refer to objects. The important practical consequence is that assigning one name to another does not automatically create an independent copy of a mutable object.

The type of a value determines what operations make sense. You can add two integers, concatenate two strings, compare booleans, or multiply a number by a number. If a source system gives you text, Python will not silently guess that the text should participate in numeric calculations. You must make that conversion deliberately.

For data engineering, this boundary is extremely important: external data should be treated as untrusted until you establish what it means. Parse it, convert it, validate it, and only then let the rest of the pipeline rely on those assumptions.

What you will do

In Python, a variable is simply a name you give to a piece of data. When you write price = 19.50, you are telling Python: remember this number and call it price. You can then use that name later in calculations, comparisons, or display.

Every value in Python has a type. The type tells Python what kind of data it is and what operations are allowed. Text like "hello" is a string (str). Whole numbers like 4 are integers (int). Numbers with decimals like 19.50 are floating-point numbers (float). True or False values are booleans (bool). And None represents the absence of a value, meaning nothing is stored.

A type hint is an optional annotation you add after a variable name, like price: float = 19.50. Type hints do not change how your program runs. They serve as documentation for humans and code editors, making your intent clear. Casting means converting a value from one type to another, such as float("19.50") which turns the text "19.50" into the number 19.50.

Why this skill

Imagine an e-commerce company receives a CSV file of daily orders from a payment processor. Every value in a CSV file arrives as text, even numbers. The price column contains "19.50" (a string), not 19.50 (a number). If a data engineer tries to multiply "19.50" by 4 without converting it first, Python will not produce 78.0. Instead, it repeats the string four times: "19.5019.5019.5019.50". The invoice total becomes garbage, customers get wrong bills, and the finance team loses trust in the data pipeline. Converting at the boundary (the moment data enters your system) prevents this entire category of bug.

How the code works

Think of variables like labeled jars on a kitchen shelf. Each jar has a label (the variable name) and contents (the value). The shape of the jar (the type) determines what you can do with the contents. You cannot pour a word into a measuring cup and expect a volume reading.

One order row stored in labeled jars
order_id → strprice_text →strqty → intis_paid → boolprice → float

Each jar holds one value. The label tells you what it means. The type tells you what operations are valid.

Python determines a value's type at runtime (when the program actually executes). You do not declare types in advance like some other languages require. However, you must explicitly convert (cast) values when the source provides them as text.

A variable can be used in an expression, which is a piece of code Python evaluates to produce a value. For example, subtotal = price * qty evaluates multiplication first and then assigns the answer to subtotal. Comparisons such as status == "paid" produce a boolean. You can combine conditions with and, or, and not, which is how a pipeline expresses rules such as "paid and not refunded".

Read an expression as a question or calculation, then store its result when you need it again.

ExpressionPlain-English readingPipeline example
price * quantitymultiply two valuescalculate a row subtotal
amount >= 0is amount at least zero?reject negative sales
status == "paid"does status equal paid?include a row in revenue
value is Noneis the value missing?separate unknown from zero
paid and not refundedare both rules true?apply a business filter

Python also has a useful idea called truthiness. In an if statement, None, False, 0, and an empty string or collection behave like false; most other values behave like true. This can make code concise, but do not use it when zero and missing mean different things. In a revenue pipeline, if amount: would incorrectly reject a legitimate amount of 0. Check amount is not None when the distinction matters.

PythonZero is not the same as missing
amount = 0.0
missing_amount = None

print(bool(amount))          # False: zero is falsy
print(bool(missing_amount))  # False: None is falsy
print(amount is None)        # False: zero is present
print(missing_amount is None)  # True: value is missing

Choose the type based on what the value means, not what it looks like.

TypeWhat it storesExample valueCommon mistake
strText or identifiers"ORD-104" or "19.50"Doing math on text that looks like a number
intWhole numbers (no decimals)4Using 0 when you mean None (missing)
floatDecimal numbers19.50Expecting exact currency math (use Decimal for money)
boolTrue or FalseTrueAssuming empty string is False in all contexts
NoneAbsence of a valueNoneConfusing None with 0 or empty string
Converting CSV text to usable numbers
CSV text arrivesfloat(price_text)int(qty_text)subtotal = price *qtyStore result

Data enters as text on the left. Each conversion step produces a typed value ready for computation.

Here is sample order data before we process it. Notice every value is a string because CSV files are plain text:

All values are strings when read from a CSV file.

order_idpriceqtystatus
ORD-10419.504paid
ORD-1058.002cancelled
ORD-10612.252paid

Worked examples

Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.

PythonBasic variable assignment and casting
order_id: str = "ORD-104"
price_text: str = "19.50"
qty: int = 4
is_paid: bool = True

price: float = float(price_text)
subtotal: float = price * qty

print(order_id, subtotal, is_paid)
print(type(price).__name__, type(subtotal).__name__)

Line by line: order_id stores a text identifier. price_text holds the raw string from the CSV. qty is already an integer. is_paid is a boolean flag. The key line is float(price_text), which converts the string "19.50" into the number 19.50. After that, price * qty produces 78.0 because both operands are numeric.

PythonProcessing a batch of orders with type conversion
orders = [
    {"order_id": "ORD-104", "price": "19.50", "qty": "4", "status": "paid"},
    {"order_id": "ORD-105", "price": "8.00", "qty": "2", "status": "cancelled"},
    {"order_id": "ORD-106", "price": "12.25", "qty": "2", "status": "paid"},
]

paid_facts: list[dict[str, object]] = []
for row in orders:
    if row["status"] == "paid":
        paid_facts.append({
            "order_id": row["order_id"],
            "subtotal": float(row["price"]) * int(row["qty"]),
        })

paid_total: float = sum(fact["subtotal"] for fact in paid_facts)
print(paid_facts)
print("paid total", paid_total)

This example processes a list of order dictionaries. For each paid order, it converts price from string to float and qty from string to int, computes the subtotal, and collects the result. The final paid_total sums all subtotals from paid orders.

PythonWrong versus right: always cast before math
# WRONG: string repetition instead of multiplication
price = "19.50"
qty = 4
print(price * qty)  # prints "19.5019.5019.5019.50" (string repeated 4 times)

# RIGHT: convert before computing
price_text = "19.50"
qty = 4
try:
    result: float = float(price_text) * qty
    print(result)  # prints 78.0
except ValueError:
    print("could not convert price:", price_text)

In the wrong version, multiplying a string by an integer repeats the string. Python does not raise an error because string repetition is a valid operation. The program runs silently with incorrect output, which is worse than a crash. In the right version, float() converts the string first, so multiplication produces the expected numeric result.

Type hints do not enforce anything at runtime

Writing price: float = "19.50" will not convert the string to a float. The hint is purely documentation. You still need to call float() explicitly to convert the value.

ValueError means the conversion failed

If the source contains garbage like "N/A" instead of a number, float("N/A") raises a ValueError. Wrap conversions in try/except at the boundary to catch and quarantine bad rows.

There is one more float surprise worth seeing before it surprises you in production: floats are not exact.

PythonFloats accumulate tiny rounding errors
total = 0.0
for _ in range(3):
    total += 0.10

print(total)          # 0.30000000000000004, not 0.3
print(round(total, 2))  # 0.3

Computers store floats in binary, and most decimal fractions (0.10 included) have no exact binary representation, so tiny rounding errors creep in. For display, round() hides it. For money that gets added up thousands of times, those tiny errors compound into real cents. That is why production billing code often uses Python's decimal.Decimal instead of float for currency, a detail worth remembering rather than using right now.

Under the hood: PyObject memory overhead

In CPython, even a simple integer is not a bare 4-byte or 8-byte CPU primitive. Every integer is a full C structure called PyObject containing an 8-byte reference count, an 8-byte pointer to the type object, and variable-length digit arrays. A single Python integer occupies 28 bytes on a 64-bit operating system. A float takes 24 bytes. When a data pipeline loads 5 million integers into a standard Python list, the overhead is not 20 megabytes; it is upwards of 140 megabytes. This explains why data engineers transition from plain Python lists to columnar formats (NumPy arrays, PyArrow buffers, Parquet files) when scaling past local batch scripts.

Real data engineering usage: Boundary parsing in banking feeds

When ingesting transaction logs from financial institutions (such as NEFT or RTGS transaction feeds from Indian banks like HDFC or ICICI), monetary values frequently arrive as strings containing commas, currency codes, or extra spaces, for example 'INR 14,500.50'. If an ingestion worker passes that raw string downstream, aggregations will either fail or silently duplicate strings. Defensive ingestion code strips whitespace, removes punctuation, converts the amount into integer paisa (1450050) or decimal.Decimal, and validates that the value is strictly positive before loading into an analytical warehouse.

Common beginner confusion: Floating-point precision in financial pipelines

A junior engineer once wrote an e-commerce commission calculation using Python floats: `total = 0.1 + 0.2`. When checking `if total == 0.3:`, the condition evaluated to False because IEEE 754 floating-point representation yields `0.30000000000000004`. When processing 500,000 orders every evening, fractions of a paisa compound into significant discrepancies, triggering reconciliation audit failures. Golden rule: store and compute money as integer cents or paisa, or use Python's decimal module.

Interview connection: 'is' vs '==' in Python

A standard junior technical interview question is: 'What is the exact difference between == and is in Python, and when must you use is?' A candidate should explain that == checks value equality (do the two objects represent the same data?), whereas is checks reference identity (do both variables point to the exact same address in memory?). In data pipelines, always write `if value is None:` rather than `if value == None:`, because None is a singleton in CPython and identity checking avoids custom object equality overloads.

Cast at the ingestion boundary

Raw files, CSVs, and API responses provide text. Clean, typed values should be created the moment data enters your pipeline, and never assumed.

Common beginner questions

Why can't Python just figure out that "19.50" is a number?

Python deliberately keeps text and numbers separate because mixing them causes subtle bugs. The string "007" is a valid piece of text but the number 7 is different. Explicit conversion forces you to decide what the data means.

What is the difference between int and float?

An int stores whole numbers without decimals, like 4, 100, or 0. A float stores numbers with decimal points, like 19.50 or 3.14. Use int for counts and quantities. Use float for measured amounts and prices.

What is the difference between = and ==?

A single = assigns a value to a name: price = 19.50 means "store 19.50 under the name price." A double == compares two values and gives back True or False: price == 19.50 asks "does price currently equal 19.50?" Using = where you meant == is one of the most common beginner typos, and Python's error messages usually point straight at it.

When should I use None?

Use None when a value is genuinely missing or unknown. For example, a customer who has not provided a phone number should have phone = None, not phone = "" or phone = 0. None makes the absence explicit.

What comes next

Now that you can store and convert individual values, the next lesson teaches you how to group related values together using lists, dictionaries, and other data structures.

Practice

Run Sample to see variable names, values, and their runtime types printed. Then complete the Exercise: cast the supplied price text to a float, multiply by qty, store the answer in a variable called result with a float type hint, and print 78.0.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Knowledge Check

Confirm your understanding of the concepts taught in this lesson.

Question 1 of 1

Why is using the 'with' statement (context manager) essential when writing ETL output to files or managing database connections?

Rate:
Was this useful?
Why Python for data engineers?DE data structures