reduceByKey combines values for each key on the map side, before the shuffle. groupByKey sends every value across the network first and groups them after. So reduceByKey shuffles far less data and cannot blow up memory on one key the way groupByKey can.
Example: sum per key
pairs = sc.parallelize([("a", 1), ("b", 1), ("a", 1), ("a", 1), ("b", 1)])
pairs.reduceByKey(lambda x, y: x + y) # good
pairs.groupByKey().mapValues(sum) # same result, more workSay partition 1 holds ("a",1), ("a",1), ("b",1). With reduceByKey, that partition first reduces locally to ("a",2), ("b",1), and only those two records enter the shuffle. With groupByKey, all three records are shuffled. With 500 million records and a few thousand keys, the difference is a few thousand records per partition versus the entire dataset.
The memory problem
groupByKey has to put all the values of one key into a single iterable on one executor. If one key is hot, say 200 million rows for user_id = null, that task runs out of memory. reduceByKey never holds all values, only the running result.
In the DataFrame API
You usually do not choose. df.groupBy("k").agg(F.sum("v")) already does a partial aggregation before the shuffle, then a final one after. That is the same idea built in. So the advice matters mostly when you write RDD code or use collect_list, which does bring all values together.
When the output type differs
reduceByKey needs the output to have the same type as the input values. If you want something else, such as a (sum, count) pair from plain numbers, use aggregateByKey or combineByKey, which take a start value and two functions: one to merge a value into the running result, and one to merge two partial results.