Interviewers want a clear path from sources to consumers, plus quality, orchestration, and failure thinking, not a random tool dump.
End-to-end architecture
Sources
Postgres (orders, users) Stripe API Clickstream
| | |
CDC (Debezium) batch pull Kafka
\ | /
\______________________|_______________/
|
Kafka / landing
|
Object storage RAW (bronze)
JSON/Parquet, append-only
|
Spark / dbt / SQL
silver: cleaned, deduped, conformed
gold: fct_orders, dim_user, revenue marts
|
+------------------+------------------+
v v v
Looker/BI ML features Ops alertsHow to talk through it (checklist)
1. Requirements: freshness (hourly BI vs realtime fraud), volume, PII. 2. Ingestion: CDC for DB tables; scheduled API pull for Stripe; Kafka for events. 3. Storage: raw bronze immutable; silver cleaned; gold marts (medallion). 4. Transforms: Spark for heavy joins; dbt for warehouse SQL models/tests. 5. Orchestration: Airflow DAG with sensors, retries, SLAs. 6. Quality: uniqueness on order_id, not-null amounts, reconciliation vs source counts. 7. Serving: BI on gold; never point dashboards at raw. 8. Ops: monitors on freshness, failed tasks, Kafka lag; idempotent partition loads.
Tiny grain examples
fct_orders: one row per orderfct_order_items: one row per line itemdim_user: SCD2 if city history matters
Interview tip: Draw sources → land → transform → serve, then spend time on idempotency, tests, and SLAs. Tools are secondary to the data path.