Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is ORC?

File formats & storage · Formats & Table Formats

What is ORC?

Easyformat-04
ORCHivecolumnarParquet

Question

What is Apache ORC, and when would you use it?

Solution

ORC (Optimized Row Columnar) is a columnar file format originally built for the Hive ecosystem. Like Parquet, it stores columns separately with compression and rich indexes/statistics.

Problem first. Early Hive on plain text or RCFile was slow. ORC was designed so Hive (and later Spark) could skip stripes, use bloom filters, and compress aggressively inside the Hadoop/Hive stack.

ORC traits interviewers expect

  • Columnar layout with stripes (similar idea to Parquet row groups)
  • Lightweight indexes and column statistics for predicate pushdown
  • Strong compression (often excellent space efficiency)
  • Native fit with Hive ACID tables historically
Hive / Spark SQL query
        |
   ORC reader uses stripe stats / indexes
        |
   skip stripes that cannot match filters

ORC vs Parquet (short)

Both are columnar analytics formats. Parquet is the more common default in modern multi-engine lakes (Spark + Trino + Iceberg/Delta). ORC remains strong in Hive-centric shops and some warehouse engines.

Interview tip: Say "ORC is Hive's classic columnar format; Parquet is the cross-engine lake default." Mention stripe stats and bloom filters if they dig deeper.

PreviousNext