Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Glue Data Catalog and crawlers

Cloud · AWS & Azure

Glue Data Catalog and crawlers

Mediumcloud-33
awsgluedata-catalogathenacrawlers

Question

What is the AWS Glue Data Catalog, and what do crawlers do?

Solution

The AWS Glue Data Catalog is a managed, Apache Hive-compatible metadata repository that stores table schemas, partition definitions, and data locations across Amazon Web Services. AWS Glue crawlers inspect files on Amazon S3 to automatically infer column types and discover partition hierarchies. Because crawlers can misinterpret schemas and generate duplicate tables during schema drift, production teams often define table structures explicitly using infrastructure code or query engines.

Central metadata for lake queries

The Glue Data Catalog acts as the single source of truth for schema definitions across the entire AWS analytical ecosystem.

S3 Storage (Parquet / Iceberg)
         |
    Glue Data Catalog (Table Schemas & Partition Metadata)
         |
    +----+-------------------+--------------------+
    |                        |                    |
Athena (SQL)           EMR (Spark)       Redshift Spectrum

Multiple analytics services rely on this catalog to parse raw storage files:

  • Glue Data Catalog replaces self-hosted Hive metastores, exposing schemas to Amazon Athena, Amazon EMR, Redshift Spectrum, and AWS Glue ETL jobs.
  • Glue crawlers scan prefixes in S3, determine file formats like Parquet or CSV, infer field data types, and register discovered partition paths into catalog tables.
  • The catalog natively supports Apache Iceberg tables, tracking metadata pointer locations so query engines access transactional lake tables without running partition discovery operations.

Crawler pitfalls in production

While crawlers simplify initial dataset onboarding, experienced data teams treat automated crawlers with caution in mission-critical pipelines:

  • Crawlers infer types based on sampled records. If an integer column receives a string or floating-point value in a later batch, the crawler may create a duplicate table or cast the column to string, breaking downstream SQL queries.
  • Inconsistent S3 directory naming can cause crawlers to create multiple fragmented tables instead of appending new partitions to an existing entity.
  • Production data platforms usually define tables explicitly using Terraform, dbt, or Athena DDL statements, reserving crawlers for raw exploratory landing zones with unpredictable source structures.
PreviousNext