Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. AWS Lambda

Cloud · AWS

AWS Lambda

Easycloud-05
awslambdaserverlesseventss3

Question

What is AWS Lambda, and when do data engineers use it?

Solution

AWS Lambda is serverless compute: you upload a function, AWS runs it on events, and you pay for invoke duration (GB-seconds). No servers to patch for that function.

Problem first. You need a small piece of glue code: when a file lands in S3, validate it and write a metadata row; or when a Kinesis batch arrives, fan out a metric. Spinning an always-on EC2 box is wasteful.

S3 ObjectCreated event
        |
        v
   Lambda function  -->  write to DynamoDB / SNS / another bucket
        |
   (scales with concurrent events, then idle = $0 compute)

Common DE use cases

  • S3-triggered light transforms / quarantine bad files
  • Kinesis / SQS consumers for small event processing
  • API backends for internal tools
  • Glue/Airflow alternatives only for *small*, short jobs (minutes, not huge Spark)

Limits you must know

  • Max runtime (minutes, not hours for giant Spark jobs)
  • Memory / CPU coupled; cold starts
  • Concurrent execution quotas
  • Not a replacement for EMR/Glue on multi-TB shuffles

Interview tip: "Lambda = event-driven serverless functions." Give one S3-trigger example, then say heavy ETL belongs on Spark/Glue/EMR, not Lambda.

PreviousNext