Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Reducing pandas memory usage

Python · pandas & Polars

Reducing pandas memory usage

Mediumpython-57
pandasmemorydtypescategoricaloptimization

Question

A DataFrame uses 12 GB of memory. How do you shrink it?

Solution

Check where the memory goes, then pick cheaper types, load fewer columns, and process in pieces. A 12 GB DataFrame is often 3 GB of real information in wasteful types.

Measure first

df.info(memory_usage="deep")
df.memory_usage(deep=True).sort_values(ascending=False)

deep=True counts the real size of strings. Text columns (Python objects) usually dominate, and the default numeric types are often wider than needed.

Shrink the types

  • Numbers: integers default to 64-bit. If age fits in 8 bits, astype("int8") uses an eighth of the memory. pd.to_numeric(col, downcast="integer") picks the smallest type that fits. Floats can often be float32.
  • Low-cardinality text: a country column with 200 distinct values repeated across 50 million rows can be a category. It stores each distinct string once, plus small integer codes.
df["country"] = df["country"].astype("category")

More type tips:

  • Strings: in pandas 2.x, Arrow-backed strings (dtype="string[pyarrow]") use much less memory than Python objects.
  • Booleans and flags stored as strings ("Y"/"N") should be real booleans.
  • Dates stored as text should be datetime64.

Load less in the first place

pd.read_csv("big.csv", usecols=["order_id", "country", "amount"],
            dtype={"country": "category", "amount": "float32"})
pd.read_parquet("big.parquet", columns=["order_id", "amount"])

Parquet lets you read only the columns, and even filter rows, before they reach memory.

Process in pieces

If it still does not fit, work in chunks (chunksize), or by partitions of the data (one month at a time), and combine small results. Delete large intermediate frames (del df_tmp) when you are done with them.

Know when to change tools

If you keep fighting memory limits, the tool may be wrong. Polars and DuckDB are more memory-efficient and can process data larger than RAM, and Spark or a warehouse handles data larger than one machine. Mention that the cheapest fix is often to avoid loading the data you do not need.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext