Amazon Kinesis is AWS's family of managed streaming services. For data engineers, Kinesis Data Streams is the closest cousin to Kafka: producers put records into a stream, consumers read them in near real time.
Clickstream / CDC / IoT
|
v
Kinesis Data Streams --> Lambda / Spark / Flink / Firehose
|
+--> S3 / Redshift / OpenSearch (via Firehose)Kinesis Data Streams basics
- Stream split into shards (throughput units; think partitions)
- Retention typically hours to days (configurable; shorter than many Kafka retention setups by default)
- Consumers can use enhanced fan-out for dedicated throughput
- Kinesis Data Firehose is a managed delivery pipe into S3/Redshift/etc. with less code
Kinesis vs Kafka
| Dimension | Kinesis Data Streams | Apache Kafka | |---|---|---| | Ops | Fully managed AWS service | Self-managed or managed (MSK / Confluent) | | Scale unit | Shard | Partition | | Ordering | Per shard | Per partition | | Ecosystem | AWS Lambda, Firehose, IAM | Broad: Connect, Streams, Flink, etc. | | Portability | AWS-locked | Multi-cloud / on-prem | | Retention | Hours–days (extendable) | Often days–weeks+ by design |
Tiny rule of thumb
- Heavy AWS + Lambda consumers + short retention → Kinesis
- Multi-cloud, long retention, rich Kafka tooling, or existing Kafka ops → Kafka / MSK
Interview tip: "Kinesis Streams ≈ managed AWS event log with shards; Kafka ≈ open streaming log with partitions." Mention Firehose as the easy land-to-S3 path.