AWS Glue job bookmarks persist pipeline execution state across runs, tracking which Amazon S3 objects or JDBC records have already been processed so subsequent job executions ingest only new data. Engineers can enable, pause, or rewind bookmarks depending on whether they are running scheduled batches, debugging code, or backfilling historical data. However, bookmarks rely on immutable files and primary keys, meaning in-place file modifications break tracking and downstream writers still require idempotent write logic.
State tracking mechanisms
Job bookmarks maintain a persistent ledger of processed source data across successive executions of an AWS Glue script.
Run 1: Ingests files A, B, C ---> Bookmarks mark A, B, C processed Run 2: S3 has A, B, C, D, E ---> Bookmarks filter out A, B, C ---> Ingests D, E
Engineers manage bookmark behavior using three operational states:
- Enable tells the job to record source state at the end of a successful run and skip previously processed files or rows on the next execution.
- Pause runs the job against current unread records without updating the saved state bookmark, which is ideal for testing transformations without advancing offsets.
- Reset clears the recorded history completely, causing the next run to process the entire dataset from the very beginning for a complete historical backfill.
- For S3 sources, bookmarks track object keys, file sizes, and modification timestamps. For JDBC database sources, bookmarks track sequential primary keys or monotonic update timestamps.
Failure modes and idempotency
Bookmarks simplify incremental ETL but introduce operational vulnerabilities that freshers must watch out for:
- If an upstream process updates an existing S3 file in place without changing its key or timestamp properties consistently, the bookmark skips the modified record, leading to silent data loss.
- A job that crashes midway through execution does not commit its bookmark update. On restart, it reprocesses the entire batch.
- Because partial failures cause duplicate reads, downstream storage must implement idempotency. Target tables should use merge statements or partition overwrites so duplicate executions do not inflate business metrics.