1TB is routine at scale if you avoid full scans, tiny files, and unnecessary shuffles. Efficiency is about I/O, partitioning, and compute shape.
Efficient path sketch columnar storage (Parquet/ORC) + compression partition by date (read only needed days) predicate pushdown / column prune scale workers horizontally minimize shuffle (filter early, broadcast small dims) avoid millions of small files (compact)
Concrete tactics
1. Do not read all 1TB if the job only needs yesterday: partition prune. 2. Use Parquet/ORC, not CSV, for analytics scans. 3. Select only needed columns. 4. Size files sensibly (often ~128MB–1GB range as a rule of thumb). 5. Tune partitions/executors so tasks are even; fix skew. 6. Prefer incremental processing over rewriting history daily. 7. Cache only when reused multiple times in one job.
Cost mindset
Scan less data, shuffle less data, store compressed columnar. Cluster size is the last lever, not the first.
Interview tip: Start with "how much of the 1TB do I actually need to read?" Partition prune + columnar format is the fresher-to-strong answer.