Kafka Connect is a framework for moving data in and out of Kafka using reusable connectors, instead of writing custom producer/consumer apps for every integration.
Problem it solves
Without Connect, every source (MySQL, Postgres, S3, Salesforce) needs a hand-rolled ingestion service. Connect standardizes:
- Configuration
- Scaling (tasks / workers)
- Offsets / positions for connectors
- Rest API for managing connectors
- Converters (JSON, Avro, Protobuf) and Single Message Transforms (SMTs)
Postgres --(source connector)--> Kafka topics --(sink connector)--> Snowflake / S3 / ES
Architecture sketch
- Workers run connectors
- A connector splits work into tasks
- Standalone mode for demos; distributed mode for production (config stored in Kafka)
Why DE interviews love it
Most real platforms use Connect (or a managed equivalent) for CDC and landing zones. Saying "I'd use a Debezium source connector into Kafka, then a sink to the warehouse" is a strong architecture answer.
Interview tip: "Connect = integration runtime for Kafka. Prefer connectors over custom glue when a maintained connector exists."