Spark has to turn objects into bytes whenever it moves them between machines or stores them in a compact form. That covers shuffles, caching with a serialized storage level, broadcast values and the closures it sends to tasks. Java serialization is the default for RDD objects. Kryo is a faster, more compact alternative.
Java vs Kryo
Java's built-in serialization is flexible but slow and writes bigger output, because it includes class metadata for every object. Kryo writes much smaller output and is often several times faster. You switch it on with:
spark = (SparkSession.builder
.config("spark.serializer", "org.apache.spark.serializer.KryoSerializer")
.getOrCreate())For best results in Scala or Java, you also register your classes, so Kryo writes a small id instead of the full class name. Without registration it still works, but is less compact.
Does it matter for DataFrames
Much less. DataFrame and Dataset operations use Tungsten's own binary row format with encoders, not Java or Kryo serialization for the rows. Kryo mostly matters if you use RDDs of your own JVM objects, or datasets of custom classes.
What it means in PySpark
PySpark adds a second layer. Python objects are pickled to move between the Python worker and the JVM. Row-by-row pickling is one of the main costs of Python UDFs and RDD lambdas. Pandas UDFs reduce it by sending columnar batches through Apache Arrow instead.
JVM executor <--pickle or Arrow--> Python worker process
What to say
For a DataFrame-based PySpark job, say serialization is rarely the first thing to tune, because the JVM side uses Tungsten and the Python side is cheaper if you avoid UDFs. For RDD-heavy Scala code, Kryo is a quick, safe win.