Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep

Interview prep

Data engineering interview questions

Theory questions across SQL, Python, Spark, Airflow, dbt, Kafka, batch & streaming, data modeling, cloud, file formats, data quality, pipelines, and behavioral interviews. Each one has a plain explanation, examples, and interview shortcuts. Filter by topic below, open any question for the full solution, browse recommended books and official documentation, or try a random set in a mock interview.

Mock interviewsCoding problems

SQL interview questions (40)

  • INNER JOIN vs LEFT JOIN
  • WHERE vs HAVING
  • What is a self-join?
  • UNION vs UNION ALL
  • Order of SQL clause execution
  • What is CROSS JOIN?
  • ON vs WHERE in a JOIN
  • What is an anti-join?
  • What is a semi-join?
  • What is NATURAL JOIN?
  • COUNT(*) vs COUNT(column)
  • What is GROUP BY?
  • SUM(CASE WHEN...) and conditional aggregation
  • What is ROLLUP?
  • What is CUBE?
  • What is a window function?
  • What does OVER() do? What is PARTITION BY?
  • ROW_NUMBER vs RANK vs DENSE_RANK
  • LAG and LEAD
  • What is a running total?
  • ROWS vs RANGE
  • FIRST_VALUE and LAST_VALUE
  • What is NTILE()?
  • PERCENT_RANK vs CUME_DIST
  • Can window functions be used in WHERE?
  • What is NULL?
  • COALESCE vs NULLIF
  • What happens when NULL is added to a number?
  • What is three-valued logic?
  • CHAR vs VARCHAR vs TEXT
  • What is an execution plan?
  • What is an index?
  • What is a covering index?
  • What is predicate pushdown?
  • Clustered vs non-clustered index
  • What is a primary key?
  • Primary key vs foreign key?
  • What is normalization?
  • DELETE vs TRUNCATE vs DROP?
  • What is a CTE?

Python interview questions (40)

  • List vs tuple
  • Set vs dictionary
  • List comprehension
  • == vs is
  • append vs extend
  • Dictionary comprehension
  • del vs remove vs pop
  • What is a generator?
  • What is yield?
  • What is an iterator?
  • What is a decorator?
  • *args vs **kwargs
  • Lambda function
  • Function vs method
  • What is a closure?
  • return vs yield
  • What are type hints?
  • What is functools.wraps?
  • What is a higher-order function?
  • Shallow copy vs deep copy
  • Class vs object
  • What is inheritance?
  • What is polymorphism?
  • What is encapsulation?
  • Instance vs class vs static method
  • What is the GIL?
  • multiprocessing vs threading
  • What is a context manager?
  • __enter__() and __exit__()
  • __str__() vs __repr__()
  • How does Python manage memory?
  • sys.getsizeof() vs actual memory usage
  • What is interning?
  • range() in Python 2 vs Python 3
  • What are __slots__?
  • Mutable vs immutable objects
  • None vs False vs 0
  • What is exception handling?
  • What is a virtual environment?
  • Difference between copy() and assignment

PySpark interview questions (40)

  • What is PySpark?
  • Spark architecture (Driver, Cluster Manager, Executors)
  • What is SparkSession?
  • RDD vs DataFrame vs Dataset
  • What is lazy evaluation?
  • Transformations vs actions
  • What is a DAG? What is Catalyst?
  • Narrow vs wide transformations
  • What is a shuffle?
  • What is Spark UI?
  • Read CSV with explicit schema
  • select() vs withColumn()
  • Handling NULLs in PySpark
  • filter() vs where()
  • How do you explode an array?
  • dropDuplicates vs distinct
  • How do you pivot a DataFrame?
  • coalesce vs repartition
  • Write partitioned Parquet
  • overwrite vs append
  • PySpark join types (incl left anti/semi)
  • What is a broadcast join?
  • What is data skew?
  • How do you fix data skew?
  • What is AQE?
  • cache vs persist
  • When should you cache?
  • What is a checkpoint?
  • What is the small-files problem?
  • How do you optimize a slow PySpark job?
  • What is a UDF?
  • What is a Pandas UDF?
  • What is Structured Streaming?
  • What is a watermark?
  • What is exactly-once processing?
  • repartition() vs partitionBy()
  • What is partition pruning?
  • What is predicate pushdown in Spark?
  • What is an executor?
  • What is a Spark stage?

Airflow & DAGs interview questions (38)

  • What is Apache Airflow, and what problem does it solve?
  • What does DAG stand for?
  • Core components of Airflow architecture
  • Operator vs Task vs Task Instance
  • How do you define task dependencies?
  • What are Sensors?
  • What is an Airflow Hook?
  • What is idempotency?
  • What is the Metadata Database?
  • What is the Airflow UI?
  • SequentialExecutor vs LocalExecutor vs CeleryExecutor
  • KubernetesExecutor vs CeleryExecutor
  • What is the role of an Airflow Worker?
  • What is DAG Serialization?
  • How do you scale Airflow to thousands of DAGs?
  • What is a Pool?
  • What happens if the Scheduler dies?
  • What are Zombie and Undead tasks?
  • How does Airflow handle logs?
  • What is the Triggerer?
  • What is XCom?
  • What is TaskFlow API?
  • What is catchup?
  • logical_date vs actual runtime
  • Passing parameters on manual trigger
  • What is Jinja templating?
  • What are Trigger Rules?
  • How do you handle branching?
  • TaskGroups vs SubDAGs
  • How do you handle secrets?
  • How do you unit test a DAG?
  • Task is up_for_retry: what does it mean?
  • Works locally but fails in Airflow: why?
  • What is depends_on_past?
  • How do you backfill six months?
  • What is a DAG Run?
  • Retry vs rerun
  • What is start_date?

dbt interview questions (38)

  • What is dbt?
  • What is a dbt model?
  • Four materializations
  • What is ref()?
  • source() vs ref()
  • Where to configure materialization
  • What is the dbt DAG?
  • What is dbt_project.yml?
  • What is profiles.yml?
  • dbt run vs dbt build
  • Schema tests: unique, not_null, accepted_values, relationships
  • Schema tests vs data tests
  • Custom generic test
  • Test severity: error vs warn
  • dbt test --select
  • Source freshness
  • dbt docs generate and serve
  • What are exposures?
  • Incremental model flow
  • What is is_incremental()?
  • Four incremental strategies
  • Incremental without unique_key
  • What is --full-refresh?
  • insert_overwrite vs merge
  • on_schema_change
  • Late-arriving data and lookback
  • Snapshots vs incremental models
  • What is {{ this }}?
  • Jinja in dbt
  • What is a dbt macro?
  • run_query() risks
  • Packages and dbt deps
  • dbt_utils highlights
  • pre-hooks and post-hooks
  • Secrets with env_var
  • What is staging in dbt?
  • Lineage in dbt
  • What are dbt seeds?

Kafka interview questions (33)

  • Apache Kafka vs RabbitMQ
  • Producer, Broker, Consumer, Topic, Partition
  • What is a Consumer Group?
  • Topic vs Partition and ordering guarantees
  • What is a Kafka broker?
  • ZooKeeper vs KRaft
  • What is the Kafka Controller?
  • Replicas: leader and follower
  • What is ISR (In-Sync Replicas)?
  • Kafka vs cloud Pub/Sub (e.g. Google Pub/Sub)
  • Producer write flow
  • Producer acks: 0, 1, and all
  • Idempotent producer
  • At-most-once, at-least-once, exactly-once
  • Consumers and offsets
  • auto.offset.reset: earliest vs latest
  • What is consumer lag?
  • What is a consumer rebalance?
  • Eager vs Cooperative rebalancing
  • Static membership
  • What is log compaction?
  • Delete vs compact retention
  • What is Schema Registry?
  • Schema compatibility: backward, forward, full
  • What is Kafka Streams?
  • What is ksqlDB?
  • What is Kafka Connect?
  • Source vs Sink connectors
  • Exactly-once in Kafka Streams
  • Kafka vs Amazon Kinesis
  • What exactly is an offset?
  • Why do we use a partition key?
  • Why can't you scale consumers infinitely in a group?

Batch & Streaming interview questions (33)

  • Batch vs stream
  • Event time vs processing time
  • What is watermarking?
  • Micro-batch vs true streaming
  • Stateful vs stateless processing
  • Tumbling, sliding, and session windows
  • Keyed vs non-keyed streams
  • What is backpressure?
  • At-least-once vs exactly-once
  • What is checkpointing?
  • What is MapReduce?
  • Hadoop MapReduce vs Spark
  • RDD vs DataFrame vs Dataset
  • What is lazy evaluation in Spark?
  • Transformations vs actions
  • Narrow vs wide transformations
  • What is a shuffle?
  • Cache vs persist
  • Repartition vs coalesce
  • The small files problem
  • Flink vs Spark Streaming
  • Structured Streaming vs DStreams
  • Event-driven vs request-response
  • What is CQRS?
  • What is event sourcing?
  • Lambda architecture
  • What is Kappa architecture?
  • What is CDC?
  • Polling vs push
  • OLTP vs OLAP
  • What is a data pipeline?
  • ETL vs ELT
  • Exactly-once vs idempotency

Data modeling interview questions (35)

  • OLTP vs OLAP
  • Normalized vs denormalized
  • What is a star schema?
  • What is a snowflake schema?
  • Fact table and three fact types
  • What is a dimension table?
  • What is grain?
  • SCD overview: Type 1, 2, and 3
  • When to use SCD Type 1 vs Type 2
  • How to implement SCD Type 2
  • What is a surrogate key?
  • What is a composite key?
  • Primary key vs unique constraint
  • What is a foreign key?
  • Data Vault: hubs, links, and satellites
  • Data Vault vs star schema
  • Conformed vs local dimension
  • Why conformed dimensions matter
  • Many-to-many modeling
  • What is a bridge table?
  • Transactional fact vs periodic snapshot
  • Late-arriving dimensions
  • Dimensional modeling vs ER modeling
  • How to decide grain
  • Design an e-commerce star schema
  • What is a degenerate dimension?
  • What is a role-playing dimension?
  • What is a junk dimension?
  • What is a factless fact table?
  • What is cardinality?
  • What is a natural key?
  • Primary key of a fact table
  • Role of the date dimension
  • What is a semantic layer?
  • Grain vs granularity

Cloud interview questions (25)

  • Amazon S3 and storage classes
  • AWS Glue vs Amazon EMR
  • Amazon Redshift (and vs Snowflake)
  • Amazon Kinesis (and vs Kafka)
  • AWS Lambda
  • Google Cloud Storage (and vs S3)
  • BigQuery
  • Cloud Dataflow and Apache Beam
  • Cloud Pub/Sub (and vs Kafka)
  • Cloud Composer
  • ADLS Gen2
  • Azure Synapse (and vs Snowflake)
  • ADF vs Airflow
  • Azure Databricks
  • Azure Event Hubs
  • Serverless vs provisioned
  • Infrastructure as Code (Terraform, CloudFormation, Pulumi)
  • Horizontal vs vertical scaling
  • Managed vs self-managed
  • Optimize cloud costs (5 strategies)
  • What is IAM?
  • What is a VPC?
  • Object storage vs block storage
  • What is a data lake?
  • Data lake vs data warehouse

File formats & storage interview questions (20)

  • Row vs columnar storage
  • What is Parquet?
  • What is Avro? Avro vs Parquet
  • What is ORC?
  • What is Delta Lake?
  • What is Iceberg? Iceberg vs Delta
  • What is Apache Hudi? Compare Delta / Iceberg / Hudi
  • What is predicate pushdown?
  • What is column pruning?
  • Partitioning in file storage
  • The small files problem
  • What is Z-ordering?
  • What is data skipping?
  • Compression vs encoding
  • Schema evolution
  • Schema-on-read vs schema-on-write
  • File compaction
  • What is partition pruning?
  • What is clustering?
  • Choosing a good partition column

Data quality interview questions (20)

  • Data quality and the six dimensions
  • Validation vs verification
  • Common data quality tests
  • What is Great Expectations?
  • dbt tests for data quality
  • What is data profiling?
  • What is data lineage?
  • What is a data catalog?
  • Observability vs monitoring
  • Anomaly detection for data
  • What is a data contract?
  • What is schema drift?
  • Data freshness
  • Data quality SLA
  • Circuit breaker for data pipelines
  • Dimension vs metric in data quality
  • What is a quality gate?
  • Data reconciliation
  • SLA vs SLO vs SLI
  • Root-cause analysis for data incidents

Behavioral interview questions (28)

  • Tell me about yourself (2-minute pitch)
  • Mistake in a pipeline (STAR)
  • Improved a slow or failing pipeline
  • Disagreement with a stakeholder
  • Most complex pipeline you built
  • Prioritize urgent requests
  • Learned a technology quickly
  • How you ensure data quality
  • Debug a production data issue
  • How you document pipelines
  • Optimized for cost or performance
  • Competing priorities
  • Push back on a stakeholder
  • Mentoring juniors or peers
  • Why this company?
  • Why should we hire you?
  • Biggest weakness
  • Conflict with a teammate
  • A time you failed
  • Where do you see yourself in 3–5 years?
  • When you don't know the answer
  • Working under tight deadlines
  • Handling repetitive work
  • What motivates you?
  • Independent work vs teamwork
  • Receiving code review feedback
  • Handling ambiguity
  • Questions to ask the interviewer

Pipelines & scenarios interview questions (30)

  • ETL vs ELT
  • What is a data pipeline?
  • What is batch processing?
  • What is streaming processing?
  • What is CDC?
  • What is idempotency?
  • What is incremental processing?
  • What is deduplication?
  • What is data partitioning?
  • What is data skew?
  • What is a shuffle?
  • What is fault tolerance in pipelines?
  • What is orchestration?
  • What is pipeline reliability?
  • What is a backfill?
  • What is a data warehouse?
  • What is a data lakehouse?
  • Lake vs warehouse vs lakehouse
  • How are APIs used in data engineering?
  • Design an end-to-end data pipeline
  • Design a safely retryable pipeline
  • Handle a source schema change
  • Process 1TB efficiently
  • Pipeline fails halfway
  • Reliability vs scalability vs performance
  • Horizontal vs vertical partitioning
  • What is an SLA for a pipeline?
  • How do you monitor a pipeline?
  • How do you secure a data pipeline?
  • What is least privilege?

Data platform interview questions (25)

  • ETL vs ELT
  • Batch vs streaming
  • What is idempotency?
  • What is a backfill?
  • Late data
  • Pipeline failure handling
  • Data lake vs data warehouse
  • Medallion architecture
  • What is a star schema?
  • SCD Type 1 vs Type 2
  • Grain of a table
  • Normalized vs denormalized
  • Data quality tests
  • Freshness SLA
  • Schema evolution
  • What are data contracts?
  • Observability for data
  • What is CDC?
  • Orchestration vs transformation
  • OLTP vs OLAP
  • Partitioning in analytics
  • Cost awareness for data platforms
  • File formats: Parquet vs CSV
  • Batch + streaming together
  • End-to-end mental model
LakeBenchPractice today. Build tomorrow.

Warehouse practice that runs in the tab, not on a cluster. Learn concepts, solve interview drills, and mock the round in one place.

Product

  • Studio sandbox
  • Capstone projects

Practice

  • SQL interview questions
  • PySpark interview questions
  • Python interview questions
  • DE theory questions
  • LeetCode for data engineers

Company

  • About
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds & cancellation
  • Shipping & delivery

© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.

No cluster. No install. Just the tab.