The trigger controls how often Structured Streaming starts a new micro-batch. The options are the default (as fast as possible), a fixed interval, available-now, and continuous.
The options
# default: start the next batch as soon as the last one finishes df.writeStream.start() # fixed interval df.writeStream.trigger(processingTime="1 minute").start() # process everything available, then stop df.writeStream.trigger(availableNow=True).start() # continuous, experimental df.writeStream.trigger(continuous="1 second").start()
Default
With no trigger, Spark starts a new batch right after the previous one ends, with whatever data arrived. This gives low latency but can create many small batches, and small files in the sink, when data arrives slowly.
processingTime
Start a batch every N seconds or minutes. If a batch takes longer than the interval, the next one starts right after it. Choose the interval to match how fresh the data needs to be. Writing every minute instead of every few seconds gives larger files and fewer commits, and lower cost.
availableNow
Processes all the data available at the moment, possibly in several batches that respect rate limits, and then stops the query. It is the recommended replacement for the older once trigger, which processed everything in a single batch and could run out of memory on a big backlog. It lets you use streaming semantics (checkpointed progress, exactly-once handling of files and offsets) on a schedule, such as an hourly job started by Airflow. The cluster only runs while there is work, so cost is much lower than keeping a stream up all day, and you still get incremental processing.
continuous
Processes records as they arrive with latencies of about a millisecond, instead of micro-batches. It is experimental, supports only some operations (maps and filters, not aggregations) and some sources and sinks, and gives at-least-once guarantees. Few production jobs use it.
How to choose
Ask what freshness the business needs. If minutes are fine, processingTime or availableNow on a schedule is cheaper and simpler than a stream that never stops. Needing seconds justifies an always-on query. Say that cost goes up quickly as latency goes down.