Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. explode vs explode_outer vs posexplode

PySpark · DataFrame API in Practice

explode vs explode_outer vs posexplode

Easypyspark-76
explodeexplode-outerposexplodearrays

Question

What is the difference between explode, explode_outer and posexplode?

Solution

All three turn an array column into rows. They differ in what they do with empty arrays and whether they give you the position.

Side by side

Take this table:

order_id  items
1         [pen, book]
2         []
3         NULL
df.select("order_id", F.explode("items").alias("item"))
# 1 pen
# 1 book
# (orders 2 and 3 are gone)

df.select("order_id", F.explode_outer("items").alias("item"))
# 1 pen
# 1 book
# 2 NULL
# 3 NULL

df.select("order_id", F.posexplode("items").alias("pos", "item"))
# 1 0 pen
# 1 1 book

In short:

  • explode drops rows whose array is empty or NULL. Those orders vanish from the result.
  • explode_outer keeps them, with NULL in the exploded column.
  • posexplode also returns the index (starting at 0) of each element. There is a posexplode_outer as well.

Which to use

If you only care about the items that exist, explode is fine. If you need to count orders, or to report "orders with no items", it is a bug to use plain explode, because it quietly removes the empty ones. Check the row count after the explode.

Use posexplode when order matters. For example, an array of page views where the position is the step in the journey, or an array of coordinates. Arrays do not store an index of their own in the result, so without posexplode you lose the order information once rows are processed in parallel.

Maps

explode on a map column gives two columns, key and value, one row per entry. That is a handy way to turn a map of attributes into a long table.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext