Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Distinct Count

PySpark data engineering interview problem. Difficulty: beginner. Pattern: Aggregation. About 12 minutes. Free to practice.

Count distinct user_id values into a one-row frame. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

From df, compute the number of distinct user_id values. Return a one-row DataFrame with column distinct_users. Assign result. Prefer distinct().count() over collecting rows.

Constraints

  • One output row.
  • Column name distinct_users.

Examples

Input: events fixture Output: distinct_users 3 Users 1, 2, and 3 appear in the fixture.

Topics: lakebench, pyspark, distinct, count.

More PySpark interview questions · All interview problems · Learn data engineering

beginner

Distinct Count

Interview-style drill: Count distinct user_id values into a one-row frame.

From `df`, compute the number of distinct `user_id` values. Return a one-row DataFrame with column `distinct_users`. Assign `result`. Prefer distinct().count() over collecting rows.