Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Triggers in Structured Streaming

PySpark · Streaming & Newer Spark

Triggers in Structured Streaming

Hardpyspark-85
structured-streamingtriggersavailablenowcost

Question

What trigger options exist in Structured Streaming?

Solution

The trigger controls how often Structured Streaming starts a new micro-batch. The options are the default (as fast as possible), a fixed interval, available-now, and continuous.

The options

# default: start the next batch as soon as the last one finishes
df.writeStream.start()

# fixed interval
df.writeStream.trigger(processingTime="1 minute").start()

# process everything available, then stop
df.writeStream.trigger(availableNow=True).start()

# continuous, experimental
df.writeStream.trigger(continuous="1 second").start()

Default

With no trigger, Spark starts a new batch right after the previous one ends, with whatever data arrived. This gives low latency but can create many small batches, and small files in the sink, when data arrives slowly.

processingTime

Start a batch every N seconds or minutes. If a batch takes longer than the interval, the next one starts right after it. Choose the interval to match how fresh the data needs to be. Writing every minute instead of every few seconds gives larger files and fewer commits, and lower cost.

availableNow

Processes all the data available at the moment, possibly in several batches that respect rate limits, and then stops the query. It is the recommended replacement for the older once trigger, which processed everything in a single batch and could run out of memory on a big backlog. It lets you use streaming semantics (checkpointed progress, exactly-once handling of files and offsets) on a schedule, such as an hourly job started by Airflow. The cluster only runs while there is work, so cost is much lower than keeping a stream up all day, and you still get incremental processing.

continuous

Processes records as they arrive with latencies of about a millisecond, instead of micro-batches. It is experimental, supports only some operations (maps and filters, not aggregations) and some sources and sinks, and gives at-least-once guarantees. Few production jobs use it.

How to choose

Ask what freshness the business needs. If minutes are fine, processingTime or availableNow on a schedule is cheaper and simpler than a stream that never stops. Needing seconds justifies an always-on query. Say that cost goes up quickly as latency goes down.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext