union combines two DataFrames by column position, and unionByName combines them by column name. The first one can silently put values into the wrong columns if the order differs.
The silent mistake
a = spark.createDataFrame([(1, "IN")], ["id", "country"])
b = spark.createDataFrame([("US", 2)], ["country", "id"])
a.union(b).show()
# id country
# 1 IN
# US 2 <- "US" landed in id, 2 landed in countrySpark only checks that the number of columns matches (and that types can be reconciled). It does not look at names. If types are compatible, there is no error at all. Here the id column became a string, with "US" in it.
a.unionByName(b).show() # id country # 1 IN # 2 US
unionByName lines up id with id and country with country, whatever the order.
Missing columns
If one side has an extra column, unionByName fails by default. Pass allowMissingColumns=True and the missing values are filled with NULL:
a.unionByName(b_with_extra, allowMissingColumns=True)
That is useful when combining files from different dates after a column was added.
No dedupe
Neither one removes duplicates. Both behave like SQL UNION ALL. If you want UNION semantics, add .distinct() or .dropDuplicates([...]) afterwards. Deduplication needs a shuffle, so only pay for it when you need it.
Good practice
Prefer unionByName in pipelines. A schema change upstream that reorders columns will not corrupt your data. Even better, select the columns in an explicit order before the union, so the intent is visible in the code, and compare schemas when combining files you do not control.