Check where the memory goes, then pick cheaper types, load fewer columns, and process in pieces. A 12 GB DataFrame is often 3 GB of real information in wasteful types.
Measure first
df.info(memory_usage="deep") df.memory_usage(deep=True).sort_values(ascending=False)
deep=True counts the real size of strings. Text columns (Python objects) usually dominate, and the default numeric types are often wider than needed.
Shrink the types
- Numbers: integers default to 64-bit. If
agefits in 8 bits,astype("int8")uses an eighth of the memory.pd.to_numeric(col, downcast="integer")picks the smallest type that fits. Floats can often befloat32. - Low-cardinality text: a
countrycolumn with 200 distinct values repeated across 50 million rows can be a category. It stores each distinct string once, plus small integer codes.
df["country"] = df["country"].astype("category")More type tips:
- Strings: in pandas 2.x, Arrow-backed strings (
dtype="string[pyarrow]") use much less memory than Python objects. - Booleans and flags stored as strings ("Y"/"N") should be real booleans.
- Dates stored as text should be
datetime64.
Load less in the first place
pd.read_csv("big.csv", usecols=["order_id", "country", "amount"],
dtype={"country": "category", "amount": "float32"})
pd.read_parquet("big.parquet", columns=["order_id", "amount"])Parquet lets you read only the columns, and even filter rows, before they reach memory.
Process in pieces
If it still does not fit, work in chunks (chunksize), or by partitions of the data (one month at a time), and combine small results. Delete large intermediate frames (del df_tmp) when you are done with them.
Know when to change tools
If you keep fighting memory limits, the tool may be wrong. Polars and DuckDB are more memory-efficient and can process data larger than RAM, and Spark or a warehouse handles data larger than one machine. Mention that the cheapest fix is often to avoid loading the data you do not need.