Streaming data directly into cloud data warehouses and lakehouses is accomplished using specialized ingestion mechanisms: the Storage Write API and Pub/Sub direct subscriptions in BigQuery, Snowpipe Streaming in Snowflake, and Auto Loader with Structured Streaming into Delta Lake on Databricks. These solutions make incoming events queryable within seconds, but continuous writes create small-file fragmentation and incur higher cloud expenses compared to scheduled micro-batching. Data engineers must weigh the business need for sub-minute freshness against background compaction overhead and ingestion pricing.
Ingestion mechanisms across major platforms
Each major cloud platform provides purpose-built ingestion paths optimized for high throughput:
- Google BigQuery: The BigQuery Storage Write API offers low-latency streaming ingestion with exactly-once delivery guarantees using committed streams with explicit stream offsets. Alternatively, serverless BigQuery subscriptions in Google Cloud Pub/Sub write messages directly from topics into BigQuery tables without requiring any intermediate compute infrastructure.
- Snowflake: Snowpipe Streaming bypasses staged files and warehouse virtual compute, writing rows directly from client applications into Snowflake micro-partitions via low-latency gRPC channels, reducing latency to seconds.
- Databricks Delta Lake: Spark Structured Streaming pairs with Auto Loader to continuously ingest streaming data from Kafka or cloud object storage. Ingestion writes directly into Delta Lake tables, using ACID transaction logs and automatic schema evolution.
The small-file dilemma and background compaction
Continuously writing streaming data into cloud object storage produces hundreds of tiny files every hour:
Continuous stream (1 record/sec) ---> Thousands of 50 KB Parquet files on S3/GCS Downstream analytical queries ---> High metadata scan overhead & slow performance
Because analytical engines achieve peak performance when reading large columnar files (between 128 MB and 512 MB), streaming pipelines require automated background maintenance.
Engineers must schedule regular compaction tasks, such as running OPTIMIZE in Delta Lake or relying on Snowflake automatic background clustering, to merge small files into consolidated blocks.
Cost trade-offs and micro-batching
True continuous streaming ingestion incurs dedicated API surcharges or demands always-on computing clusters.
If business dashboards and analytical models only require data updated every five to fifteen minutes, switching to micro-batch ingestion (such as Spark Structured Streaming with availableNow=True or scheduled Snowpipe) significantly reduces costs. Micro-batching groups records into well-sized file writes naturally, minimizing small-file creation and lowering cloud spend.