AWS Lambda is serverless compute: you upload a function, AWS runs it on events, and you pay for invoke duration (GB-seconds). No servers to patch for that function.
Problem first. You need a small piece of glue code: when a file lands in S3, validate it and write a metadata row; or when a Kinesis batch arrives, fan out a metric. Spinning an always-on EC2 box is wasteful.
S3 ObjectCreated event
|
v
Lambda function --> write to DynamoDB / SNS / another bucket
|
(scales with concurrent events, then idle = $0 compute)Common DE use cases
- S3-triggered light transforms / quarantine bad files
- Kinesis / SQS consumers for small event processing
- API backends for internal tools
- Glue/Airflow alternatives only for *small*, short jobs (minutes, not huge Spark)
Limits you must know
- Max runtime (minutes, not hours for giant Spark jobs)
- Memory / CPU coupled; cold starts
- Concurrent execution quotas
- Not a replacement for EMR/Glue on multi-TB shuffles
Interview tip: "Lambda = event-driven serverless functions." Give one S3-trigger example, then say heavy ETL belongs on Spark/Glue/EMR, not Lambda.