Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Event-driven batch

Batch & Streaming · Batch Design

Event-driven batch

Easybatch-streaming-53
event-driven-batchobject-storageeventbridgemanifests

Question

What is event-driven batch processing?

Solution

Event-driven batch processing triggers analytical batch pipelines immediately upon the arrival of data files in cloud object storage, rather than waiting for scheduled cron timers on a clock. Cloud storage notifications from Amazon S3, Google Cloud Storage, or Azure Blob Storage route events through services like EventBridge or Pub/Sub to launch orchestrator DAGs. This model eliminates idle sensor polling and shortens processing latency, but pipelines must protect against partial file uploads by using success markers or manifests while handling duplicate trigger notifications.

Replacing clock polling with storage notifications

Traditional batch scheduling launches jobs at fixed times, such as 3:00 AM every night.

If an upstream partner uploads their files early at 1:30 AM, data sits idle for ninety minutes. If the upload is delayed until 3:05 AM, the scheduled batch job fails or processes empty folders unless complex sensor polling loops are maintained.

Event-driven batching inverts this relationship:

S3 Object Upload ---> S3 Event Notification ---> AWS EventBridge / SQS ---> Airflow / Lambda

When an upstream process finishes uploading data to s3://landing-bucket/orders/, the cloud provider publishes an ObjectCreated event. The message broker triggers the processing pipeline immediately, delivering data to downstream consumers as soon as it arrives.

Guarding against partial uploads and write atomicity

A significant trap in event-driven architectures is reacting to incomplete file uploads.

When large multi-gigabyte files or distributed partition folders are uploaded, cloud storage creates objects in stages. If a notification triggers on the arrival of the first partial chunk, the downstream Spark or warehouse loader reads incomplete data, leading to corruption or parse errors.

Production pipelines prevent this through two patterns:

  • Success marker files: Configure triggers to fire only when a small _SUCCESS file is uploaded, which upstream jobs write strictly after all data parts have successfully landed.
  • Manifest files: Have upstream publishers upload a signed JSON or YAML manifest containing the list of all partition file paths, row counts, and checksums. The pipeline launches only when the manifest appears and validates all listed files before ingestion.

Deduplicating event notifications

Cloud event brokers operate with at-least-once delivery guarantees, meaning network hiccups or broker retries can generate duplicate notifications for a single file upload.

Pipelines must incorporate idempotency safeguards. Maintain a processed event log in a metadata database; when a notification arrives, verify whether its unique event ID or file checksum has already been processed before initiating downstream tasks.

PreviousNext