Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. When not to use pandas

Python · pandas & Polars

When not to use pandas

Easypython-58
pandasscalingpolarsduckdbspark

Question

When would you not use pandas for a data engineering task?

Solution

Do not use pandas when the data does not comfortably fit in one machine's memory, when you need parallel or distributed processing, or when a production pipeline needs strong guarantees that pandas does not give. For small and medium data, and for glue code, pandas is still a great choice.

Where pandas runs into trouble

  • Size: pandas holds everything in memory, and operations often make copies, so you need several times the data size in RAM. A 5 GB file can need 20 GB.
  • Speed: most operations use one core. A group-by on 200 million rows will be slow, while engines that use all cores or many machines finish much faster.
  • Typing: pandas is forgiving. A column of integers with one missing value silently becomes floats, and mixed types become object. In a production pipeline you want a schema that is enforced, not guessed.
  • Semantics: NaN and None handling, index alignment, and implicit type changes have caused many subtle bugs in jobs that run unattended.

What to use instead

Data size            Tool
up to a few GB       pandas (fine)
tens of GB, one box  Polars or DuckDB (fast, memory-efficient, can go beyond RAM)
hundreds of GB+      PySpark / Dask, or SQL in a warehouse
already in a DB      do it in SQL where the data lives

The last line matters. If the data is already in BigQuery or Snowflake, pulling it into pandas to filter and aggregate it moves the data to the work. Moving the work to the data (SQL) is almost always cheaper.

Where pandas is still the right call

Exploring a sample, plotting, small reference tables, quick scripts, notebooks, glue between systems, and the final step where a small aggregated result goes into a chart or a machine learning library that expects pandas or NumPy. Its API is also the one most analysts know.

How to answer

Say you pick the tool by size, shape of work and team skills. A good sentence: "I use pandas for small data and last-mile work, DuckDB or Polars for medium data on one machine, and Spark or the warehouse for anything big or shared." Then give a size threshold from your own experience or a tested measurement.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext