`explode` turns each array element into its own row (and optionally structs via explode on array<struct>). Use explode_outer to keep rows with null/empty arrays.
Example
from pyspark.sql import functions as F
df = spark.createDataFrame(
[(1, ["a", "b"]), (2, ["c"]), (3, None)],
["id", "tags"],
)
# Drops id=3 (null array) by default
df.select("id", F.explode("tags").alias("tag")).show()
# Keeps id=3 with tag=null
df.select("id", F.explode_outer("tags").alias("tag")).show()
# posexplode gives position
df.select("id", F.posexplode("tags").alias("pos", "tag")).show()Diagram
id | tags
1 | [a, b] explode => id | tag
2 | [c] 1 | a
1 | b
2 | cNested / map variants
explodeonmapyields key/value (useexplodecarefully; preferexplode_outer).- For JSON arrays, parse with
from_jsonthen explode.
Interview tip
Mention cardinality blow-up: exploding large arrays multiplies rows and can cause skew or OOM if unchecked.