Accepted answer
Start with the execution plan, numbers beat guesses.
Asked to compute the median of a column in SQL. Blanked on PERCENTILE_CONT syntax under pressure and tried to reconstruct it from first principles with ROW_NUMBER instead. Is that an acceptable fallback in an interview?
Accepted answer
Start with the execution plan, numbers beat guesses.
ELI5 version: it's not broken, it's just slow because it's checking way more stuff than it needs to. Narrowing what it checks is almost always the fix.
Idempotent writes with merge keys saved us during backfills.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
In our case the root cause was an implicit cast preventing pushdown.
What batch size or interval worked for you at similar scale?
Hired someone who did exactly this. Blanked on the built-in function, derived it from scratch while explaining each step. Was a stronger signal than a candidate who just knew the syntax.
Idempotent writes with merge keys saved us during backfills.
Note that merge on Delta still needs unique keys defined correctly.
Prefer a staging table plus validation gate before promoting to prod tables.
Start with the execution plan, numbers beat guesses.
Yes. Deriving median via ROW_NUMBER and COUNT, finding the middle row or rows by rank and averaging if the count is even, shows you understand what median actually computes, which is arguably more valuable to an interviewer than reciting PERCENTILE_CONT syntax from memory.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.