approx_count_distinct estimates the number of distinct values using a small fixed-size sketch, instead of tracking every distinct value. It is much cheaper than countDistinct, and it is accurate to within a few percent by default.
Why exact distinct count is expensive
To get an exact count, Spark must find all the distinct values across the whole dataset. That means shuffling the values so that equal ones meet, and deduplicating them. With 800 million distinct user ids, that is a large shuffle and a lot of memory.
How the approximate one works
It uses the HyperLogLog++ algorithm. Each partition builds a small sketch (a few kilobytes) that summarises which values it has seen. Sketches from all partitions are merged, which is cheap, and the final count is read from the merged sketch. Memory stays small however many distinct values exist, and no big shuffle of raw values is needed.
df.agg(F.approx_count_distinct("user_id", rsd=0.02).alias("users"))The rsd parameter is the maximum relative standard deviation. The default is 0.05, about 5 percent. A smaller value is more accurate and uses a larger sketch. A figure of 0.02 means the estimate is usually within 2 percent.
When to use which
- Approximate: dashboards, monitoring ("roughly how many daily active users"), data profiling, and quick checks on huge tables. A number like 1.02 million versus 1.00 million rarely changes a decision.
- Exact: anything that is billed, reconciled or audited, such as finance reports, or counting items against a legal limit. Also use exact when the data is small, since then there is little saving.
Do not mix them up
If one report uses the approximate count and another the exact count, the numbers will not match, and people will report a bug. Label the approximate ones in the output. Also, approximate counts do not add up across groups: summing the daily estimates does not give the monthly distinct count. Compute the monthly estimate from the raw data, or merge sketches if your engine exposes them.