Both Google Cloud Storage and Amazon S3 provide strong read-after-write and listing consistency, meaning readers immediately see new objects and updated metadata after a successful write. Neither service has real directory structures, treating paths instead as flat object keys with prefix strings separated by slashes. Renaming a logical folder requires copying and deleting every underlying object individually, which is why modern lakehouse formats commit changes via metadata files rather than filesystem renames.
Object storage mechanics versus POSIX filesystems
Freshers often treat cloud buckets like local filesystems, but object stores operate as distributed key-value repositories accessed through REST APIs.
Filesystem: Rename directory /staging/data to /prod/data (Atomic pointer update) Object Store: List /staging/data/ ---> Copy each object ---> Delete source objects
Understanding this physical reality prevents common pipeline bugs:
- Strong consistency is standard across both platforms. S3 adopted strong consistency for PUT, LIST, and DELETE operations in December 2020, while GCS has provided it since launch. Reading an object or listing a prefix immediately reflects completed writes.
- Virtual folders do not exist on disk. A path like orders/year=2026/file.parquet is simply a single key string. Tools like the AWS console show visual folders by splitting key names on slash delimiters.
- Renaming an entire folder is not an atomic operation. Moving a directory containing ten thousand files triggers ten thousand individual copy calls followed by ten thousand delete calls. If a worker dies halfway through, the dataset remains partially copied.
- Cloud providers enforce per-prefix request throttling limits. S3 supports 3,500 PUT and 5,500 GET requests per second per prefix. High-throughput ingestion pipelines distribute keys across diverse prefix hashes to avoid HTTP 503 throttling errors.
Why table formats replaced directory renames
Legacy Hadoop committers relied on directory renames to publish final batch outputs, which caused severe latency and partial-write corruptions on object storage:
- Modern formats like Apache Iceberg and Delta Lake write data files with unique identifiers directly into data directories.
- They commit changes by appending an atomic JSON or Avro metadata manifest that tracks active data files, bypassing directory rename operations completely.