
Community
Ask about Spark skew, SQL plans, Airflow retries, dbt tests, and interview problems. Reply in threads, vote on useful answers, and keep the conversation on the work.
can anyone explain diff btw CTE and subqueries?
I have read the docs but real-world tradeoffs are unclear. What would you optimize first? Context: Guaranteeing order per user_id across partitions Happy to share schema snippets or metrics if useful.
Calling df.toPandas() on what I estimated as a 2GB DataFrame kills the driver with an OOM. Driver has 8GB. The estimate was rows times avg_row_size, but that's clearly wrong somewhere.
Interview follow-up got me thinking, how do you actually implement this in production? Context: EMR on EKS vs traditional EMR for Spark, ops overhead? Happy to share schema snippets or metrics if useful.
Team is split on approach A vs B. Looking for experiences from similar scale. Context: Dataflow streaming with BigQuery sink, duplicate rows on retry Happy to share schema snippets or metrics if useful.
We have a Confluence space full of pipeline docs, most of which are 1-2 years stale and actively misleading at this point. What actually works to keep docs current, versus the usual "we'll keep it updated" that never happens?
Team is split on approach A vs B. Looking for experiences from similar scale. Context: SQL window function question, top 3 orders per customer Happy to share schema snippets or metrics if useful.
Loading a daily sales snapshot into the warehouse. Team is debating full partition overwrite versus merge on the natural key. Failures mid-run leave partial data with overwrite. Merge is slower but safer. What actually drives the decision for you?
I have read the docs but real-world tradeoffs are unclear. What would you optimize first? Context: Data vault hype, when is it actually justified? Happy to share schema snippets or metrics if useful.
Designing dim_customer SCD2 in Snowflake. Team is split on valid_to IS NULL for the current row versus a '9999-12-31' sentinel. Downstream dbt models and BI tools mix both patterns already. What are the real pros and cons in production pipelines?
Team is split on approach A vs B. Looking for experiences from similar scale. Context: CDC from OLTP to lakehouse, Debezium vs DMS Happy to share schema snippets or metrics if useful.
This worked in dev on sample data but fails at full volume. Details: Context: Generator pattern for reading large CSV exports from S3 Happy to share schema snippets or metrics if useful.
Students can enroll in multiple courses, courses have multiple students. Building the warehouse model, a bridge table between fact_enrollment and dim_course/dim_student, or denormalize course info directly onto the enrollment fact?
This worked in dev on sample data but fails at full volume. Details: Context: Great Expectations vs dbt tests for pipeline contracts Happy to share schema snippets or metrics if useful.
Producer added a new required field to an Avro schema. Registry accepted it under FORWARD compatibility, but an older consumer using an outdated schema failed to deserialize the new messages entirely.
Company is expanding into a second region for latency reasons. Data pipelines currently run entirely in one region. Designing the multi-region setup, active-active processing in both regions, or active-passive with failover?
Hitting a wall in prod and looking for patterns others have used. Minimal repro below but happy to share more context. Context: EMR on EKS vs traditional EMR for Spark, ops overhead? Happy to share schema snippets or metrics if useful.
Need near-real-time change capture from a vendor's MySQL database, but they won't grant binlog replication access for security reasons. Only read access to tables through a reporting replica. What's the least-bad way to approximate CDC here?
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.