Amazon Redshift relies on distribution styles to divide table rows across cluster compute slices and sort keys to order records physically within 1MB disk blocks. Choosing effective distribution keys co-locates matching join rows on the same physical slices to avoid costly network broadcast traffic, while sort keys allow execution engines to skip unneeded blocks using in-memory zone maps. Poor distribution key choices cause data skew that slows entire queries down to the speed of the slowest slice.
Distribution styles and data placement
Redshift operates as a massively parallel processing database where cluster nodes divide compute responsibilities across internal hardware slices.
KEY Distribution: Rows with matching join keys stay on the same slice ALL Distribution: Small dimension tables copied to every slice EVEN Distribution: Rows distributed round-robin when no common join key exists
Engineers select distribution strategies based on query patterns and table sizing:
- DISTSTYLE KEY distributes rows based on values in a specific column, such as customer_id. When joining two massive tables on customer_id, matching rows reside on the exact same slice, eliminating network shuffle.
- DISTSTYLE ALL copies the entire table to every node. This strategy is ideal for small, slow-changing dimension tables under a few million rows, enabling local joins against any fact table.
- DISTSTYLE EVEN spreads data uniformly across slices in a round-robin sequence, which is suitable when a table lacks clear join keys.
- DISTSTYLE AUTO delegates initial placement decisions to Redshift, transitioning from ALL to EVEN or KEY as data volume expands.
- Skewed distribution keys concentrate millions of rows onto a single slice, causing that slice to max out disk capacity and delay parallel execution while other nodes sit idle.
Block pruning and cluster maintenance
Disk organization governs how efficiently Redshift reads physical storage:
- Sort keys order data blocks on disk. Redshift stores minimum and maximum values for each 1MB block in memory zone maps, skipping irrelevant blocks during range filter scans.
- Run VACUUM to reclaim space from deleted records and re-sort disk blocks, followed by ANALYZE to update the query planner statistics.
- Modern architectures feature RA3 instances with managed storage and Redshift Serverless, which decouple compute scaling from persistent storage while retaining slice-level distribution mechanics.