Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Apache Arrow

Python · pandas & Polars

Apache Arrow

Mediumpython-61
arrowcolumnarinteroperabilityzero-copyparquet

Question

What is Apache Arrow, and why does it keep showing up in Python data tools?

Solution

Apache Arrow is a standard way to lay out tabular data in memory, in columns. Because many tools agree on the same layout, they can pass data to each other without converting or copying it. That is why you see Arrow inside pandas, Polars, DuckDB, Spark and many database drivers.

What the format is

Each column is stored as a contiguous block of memory of one type, with a separate bitmap marking nulls. A column of 1 million integers is just 1 million integers next to each other. This layout suits modern CPUs (fast scans, vectorised operations), and it is the same layout regardless of the language or library.

Why a shared format matters

Without it, moving data from one tool to another meant serialising it (to CSV, pickle, JSON) and parsing it again, which is slow and uses double the memory. With Arrow:

Polars DataFrame --(same memory layout)--> DuckDB query --> pandas DataFrame
                      zero-copy, no conversion

DuckDB can read a Polars or pandas (Arrow-backed) frame directly. Spark's Pandas UDFs and toPandas() use Arrow to move batches between the JVM and Python, which is far faster than pickling rows one by one. Arrow Flight and ADBC drivers also move data between systems in this format.

Arrow versus Parquet

They are often confused, but they have different jobs. Parquet is a file format for storing data on disk, compressed and encoded, optimised for small size and selective reads. Arrow is an in-memory format, optimised for fast computation and sharing, not for compactness. A common flow: read Parquet from disk, decode it into Arrow in memory, process it, and write Parquet again.

Practical uses for a data engineer

  • Turn on Arrow in PySpark conversions: spark.sql.execution.arrow.pyspark.enabled speeds up toPandas() and createDataFrame(pandas_df).
  • Use pyarrow to read and write Parquet and to inspect schemas.
  • Choose Arrow-backed dtypes in pandas 2.x for better memory use with strings and nulls.
  • Expect Arrow in the drivers of newer databases and warehouses, where it speeds up result transfer.

How to answer

Say "a shared columnar memory format that lets tools exchange data without copying", and contrast it with Parquet in one sentence.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext