Interview-style PySpark and Pandas problems on the same seeded warehouse used across Lakebench. Filter, aggregate, join, and window DataFrames without standing up a cluster.
65 PySpark + Pandas problems, 18 of them free. They run in your browser against the same warehouse tables used in Learn. All problems · PySpark + Pandas filter · Theory questions · Learn tracks
Writing DataFrame code that is correct and would scale: choosing the right join, using window functions for latest-record and running-total problems, avoiding collect() on large data, and explaining where a shuffle happens.
Your code runs in the tab through Pyodide with a Spark-compatible DataFrame layer, so you can practice the API and the patterns without Java, a local Spark install, or a cloud notebook.