Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Dataclasses and Pydantic for records

Python · Python for Data Pipelines

Dataclasses and Pydantic for records

Mediumpython-49
dataclassespydanticvalidationtyping

Question

When would you use a dataclass vs a Pydantic model for pipeline records?

Solution

A dataclass is a lightweight way to group fields together with no checks. Pydantic is a library that checks and converts data into the types you declare. Use a dataclass for trusted internal structures, and Pydantic where data comes in from outside.

Dataclass

from dataclasses import dataclass
from datetime import date

@dataclass(frozen=True)
class OrderKey:
    order_id: int
    order_date: date

It writes __init__, __repr__ and equality for you. But the type hints are only documentation. OrderKey(order_id="abc", order_date=5) works without complaint, and the problem appears later, far from its cause.

Pydantic

from pydantic import BaseModel, Field
from datetime import datetime

class Order(BaseModel):
    order_id: int
    amount: float = Field(ge=0)
    created_at: datetime

Order(order_id="42", amount="19.5", created_at="2025-03-01T10:00:00Z")
# order_id=42, amount=19.5, created_at=datetime(...)  -> converted
Order(order_id="abc", amount=-1, created_at="x")
# raises ValidationError listing every problem

It parses and converts values (the string "42" becomes 42), enforces constraints, and gives a clear error list. It can also export a JSON schema, which makes it a good way to write down a data contract.

When to use which

  • Pydantic at boundaries: parsing API responses, reading config, validating messages from a queue or rows from a file, and request bodies in an API. These are places where bad data actually appears.
  • Dataclasses inside the program, after validation, when you only need a tidy structure.

The cost

Validation runs Python-level (or Rust-level in Pydantic v2) work for every record. Doing it on 500 million rows in a loop is far too slow. For big data, validate with columnar tools instead: a schema in pandas, Polars or Spark, or tools like Pandera. A fair middle path is to validate a sample, or each batch's schema and key rules, and not every row object.

Mention in the answer

Say that you would use Pydantic for config and API payloads, and typed columnar schemas for bulk data.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext