Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Community

Community

Data engineering discussions

Ask about Spark skew, SQL plans, Airflow retries, dbt tests, and interview problems. Reply in threads, vote on useful answers, and keep the conversation on the work.

Ask a question
0
votes
1
answers
34
views

can anyone explain the difference between ETL and ELT with an example

Data Engineeringinterview
@nikithaSep 15, 2026134
0
votes
1
answers
12
views

why is shuffling very useful in pyspark?

Data Engineeringpysparkspark
@hunnur123ProSep 15, 2026112
1
votes
1
answers
48
views

CTE vs Subqueries

can anyone explain diff btw CTE and subqueries?

SQLsql
@hunnur123ProSep 3, 2026148
33
votes
11
answers
0
views

Guaranteeing order per user_id across partitions

I have read the docs but real-world tradeoffs are unclear. What would you optimize first? Context: Guaranteeing order per user_id across partitions Happy to share schema snippets or metrics if useful.

Kafka / Streamingkafkaordering
@harini_varmaProSep 2, 2026110
24
votes
7
answers
0
views

toPandas() OOMs the driver on a DataFrame that should easily fit

Calling df.toPandas() on what I estimated as a 2GB DataFrame kills the driver with an OOM. Driver has 8GB. The estimate was rows times avg_row_size, but that's clearly wrong somewhere.

PySparkpysparkdriveroom
@adityabhatProSep 2, 202670
38
votes
5
answers
0
views

EMR on EKS vs traditional EMR for Spark, ops overhead?

Interview follow-up got me thinking, how do you actually implement this in production? Context: EMR on EKS vs traditional EMR for Spark, ops overhead? Happy to share schema snippets or metrics if useful.

AWSawsemr
@balajinarayananSep 2, 202650
67
votes
6
answers
0
views

Dataflow streaming with BigQuery sink, duplicate rows on retry

Team is split on approach A vs B. Looking for experiences from similar scale. Context: Dataflow streaming with BigQuery sink, duplicate rows on retry Happy to share schema snippets or metrics if useful.

GCPgcpdataflow
@aaravchopraSep 2, 202660
68
votes
8
answers
0
views

How do you keep documentation from rotting on a data team?

We have a Confluence space full of pipeline docs, most of which are 1-2 years stale and actively misleading at this point. What actually works to keep docs current, versus the usual "we'll keep it updated" that never happens?

Generaldata-engineeringdocumentation
@akankshamishraSep 2, 202680
61
votes
3
answers
0
views

SQL window function question, top 3 orders per customer

Team is split on approach A vs B. Looking for experiences from similar scale. Context: SQL window function question, top 3 orders per customer Happy to share schema snippets or metrics if useful.

Interviewinterviewsql
@yaminireddyProSep 1, 202630
53
votes
8
answers
0
views

Idempotent daily ETL: overwrite vs merge, how do you choose?

Loading a daily sales snapshot into the warehouse. Team is debating full partition overwrite versus merge on the natural key. Failures mid-run leave partial data with overwrite. Merge is slower but safer. What actually drives the decision for you?

ETL / ELTetldata-qualitypython
@saranyareddySep 1, 202680
27
votes
16
answers
0
views

Data vault hype, when is it actually justified?

I have read the docs but real-world tradeoffs are unclear. What would you optimize first? Context: Data vault hype, when is it actually justified? Happy to share schema snippets or metrics if useful.

Data Modelingdata-vaultmodeling
@rajeshdesaiSep 1, 2026160
47
votes
4
answers
0
views

SCD Type 2 for customers: valid_to NULL or sentinel date?

Designing dim_customer SCD2 in Snowflake. Team is split on valid_to IS NULL for the current row versus a '9999-12-31' sentinel. Downstream dbt models and BI tools mix both patterns already. What are the real pros and cons in production pipelines?

Data Warehousingwarehousingscdmodeling
@yashwanth_sundaramSep 1, 202640
38
votes
1
answers
0
views

CDC from OLTP to lakehouse, Debezium vs DMS

Team is split on approach A vs B. Looking for experiences from similar scale. Context: CDC from OLTP to lakehouse, Debezium vs DMS Happy to share schema snippets or metrics if useful.

System Designcdcsystem-design
@prateekrathoreSep 1, 202610
41
votes
2
answers
0
views

Generator pattern for reading large CSV exports from S3

This worked in dev on sample data but fails at full volume. Details: Context: Generator pattern for reading large CSV exports from S3 Happy to share schema snippets or metrics if useful.

Pythonpythongenerators
@jyothibalakrishnanProSep 1, 202620
18
votes
3
answers
0
views

Bridge table for a many-to-many relationship, normalize or denormalize?

Students can enroll in multiple courses, courses have multiple students. Building the warehouse model, a bridge table between fact_enrollment and dim_course/dim_student, or denormalize course info directly onto the enrollment fact?

Data Modelingmodelingbridge-table
@naveenraoSep 1, 202630
1
votes
3
answers
0
views

Great Expectations vs dbt tests for pipeline contracts

This worked in dev on sample data but fails at full volume. Details: Context: Great Expectations vs dbt tests for pipeline contracts Happy to share schema snippets or metrics if useful.

ETL / ELTetlvalidation
@padmanarayananProSep 1, 202630
52
votes
6
answers
0
views

Schema Registry compatibility mode broke a downstream consumer we didn't expect

Producer added a new required field to an Avro schema. Registry accepted it under FORWARD compatibility, but an older consumer using an outdated schema failed to deserialize the new messages entirely.

Kafka / Streamingkafkaschema-registry
@rajchauhanSep 1, 202660
61
votes
2
answers
0
views

Multi-region data pipeline, active-active or active-passive?

Company is expanding into a second region for latency reasons. Data pipelines currently run entirely in one region. Designing the multi-region setup, active-active processing in both regions, or active-passive with failover?

System Designsystem-designmulti-region
@hemanaiduSep 1, 202620
24
votes
8
answers
0
views

EMR on EKS vs traditional EMR for Spark, ops overhead?

Hitting a wall in prod and looking for patterns others have used. Minimal repro below but happy to share more context. Context: EMR on EKS vs traditional EMR for Spark, ops overhead? Happy to share schema snippets or metrics if useful.

AWSawsemr
@snehaswamySep 1, 202680
83
votes
2
answers
0
views

Change data capture from MySQL without binlog access

Need near-real-time change capture from a vendor's MySQL database, but they won't grant binlog replication access for security reasons. Only read access to tables through a reporting replica. What's the least-bad way to approximate CDC here?

ETL / ELTetlcdcmysql
@nandininaiduProSep 1, 202620

Topics

Popular tags

LakeBenchPractice today. Build tomorrow.

Warehouse practice that runs in the tab, not on a cluster. Learn concepts, solve interview drills, and mock the round in one place.

Product

  • Studio sandbox
  • Capstone projects

Practice

  • SQL interview questions
  • PySpark interview questions
  • Python interview questions
  • DE theory questions
  • LeetCode for data engineers

Company

  • About
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds & cancellation
  • Shipping & delivery
  • © 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.

    No cluster. No install. Just the tab.