Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. How do you explode an array?

PySpark · DataFrame API & I/O

How do you explode an array?

Mediumpyspark-15
explodearraysdataframe-api

Question

How do you explode an array column in PySpark?

Solution

`explode` turns each array element into its own row (and optionally structs via explode on array<struct>). Use explode_outer to keep rows with null/empty arrays.

Example

from pyspark.sql import functions as F

df = spark.createDataFrame(
    [(1, ["a", "b"]), (2, ["c"]), (3, None)],
    ["id", "tags"],
)

# Drops id=3 (null array) by default
df.select("id", F.explode("tags").alias("tag")).show()

# Keeps id=3 with tag=null
df.select("id", F.explode_outer("tags").alias("tag")).show()

# posexplode gives position
df.select("id", F.posexplode("tags").alias("pos", "tag")).show()

Diagram

id | tags
1  | [a, b]     explode =>   id | tag
2  | [c]                     1  | a
                             1  | b
                             2  | c

Nested / map variants

  • explode on map yields key/value (use explode carefully; prefer explode_outer).
  • For JSON arrays, parse with from_json then explode.

Interview tip

Mention cardinality blow-up: exploding large arrays multiplies rows and can cause skew or OOM if unchecked.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext