Structured data adheres to a rigid, predefined schema organized into rows and columns, semi-structured data contains self-describing tags or key-value hierarchies without a fixed tabular schema, and unstructured data lacks any predefined conceptual data model. Structured data is stored in relational databases and warehouse tables, semi-structured formats like JSON and Avro are housed in specialized VARIANT or struct columns, and unstructured artifacts such as images, audio recordings, and PDF documents are stored as raw objects in cloud object storage. Recent advances in vector embeddings and large language models now make unstructured data searchable and queryable alongside relational datasets.
Schema rigidity and storage layouts
Understanding these three categories dictates which storage engine and processing framework to deploy:
- Structured data: Operates under a strict schema-on-write contract where every record matches defined data types. Stored in relational tables (PostgreSQL) or columnar files (Parquet) with dictionary encoding and statistics, enabling fast aggregation and strict referential integrity.
- Semi-structured data: Uses self-describing formats such as JSON, Avro, and XML that blend schema flexibility with structural organization. Modern warehouses (Snowflake, BigQuery, Databricks) store these payloads in VARIANT or JSON native column types, decomposing sub-properties into columnar micro-shreds behind the scenes to allow direct SQL path querying.
- Unstructured data: Encompasses media files, scanned contract PDFs, voice transcripts, and raw text logs that cannot fit inside tabular cells. Stored in cloud object storage like Amazon S3 or Google Cloud Storage, while tabular catalog tables track file locations, creation dates, MIME types, and access controls.
Data Type | Structural Properties | Storage Medium | Example Formats Structured | Fixed tabular schema | Warehouse tables | Parquet, ORC, SQL tables Semi-Structured | Self-describing, nested | VARIANT / Struct cols | JSON, Avro, Protocol Buffers Unstructured | Freeform binary / text | Cloud object storage | PDF, MP3, PNG, Video files
Modern AI workflows are transforming how unstructured assets are processed:
Extracting signal from unstructured blobs
Historically, data platforms treated unstructured files as opaque blobs, extracting only superficial metadata like file size and upload timestamps. To analyze the actual contents of customer call audio or legal contracts, engineers had to write cumbersome rule-based OCR and regex parsers.
The rise of generative AI and multimodal models changes this dynamic:
- Vector embeddings: Pipelines pass documents, transcripts, and audio through embedding models to generate high-dimensional vector representations.
- Vector indexing: Embeddings are indexed in vector databases or warehouse vector extensions (such as pgvector or Databricks Vector Search), enabling semantic search across unstructured corpuses.
- Relational synthesis: Large language models extract structured entities, customer sentiment scores, and summary fields from raw text, writing normalized attributes directly into downstream gold analytical tables.