Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Time-series databases

Data platform · Distributed Systems Basics

Time-series databases

Mediumdata-platform-32
time-seriesinfluxdbtimescaledbtelemetry

Question

What is a time-series database, and when would you use one?

Solution

A time-series database is a specialized storage engine optimized for timestamped, append-heavy data streams where records arrive in chronological order and queries filter predominantly over time windows. Engines like InfluxDB, TimescaleDB, and Prometheus use time-ordered physical layouts to achieve extreme write throughput, aggressive compression, automated downsampling, and TTL-based data eviction. You use a time-series database for high-frequency infrastructure metrics, financial tick data, and IoT sensor streams where low-latency write ingestion and real-time interval aggregations are required.

Timestamp ordered data structures

Time-series workloads display unique characteristics that degrade conventional relational tables:

  • Chronological append patterns: Writes arrive sequentially with recent timestamps and virtually never update historical rows, allowing engines to append to immutably sorted blocks.
  • Delta and gorilla compression: Because consecutive timestamps increment predictably, algorithms like delta-of-delta encoding compress timestamps down to fractions of a byte. Floating-point metric values compress heavily using XOR-based Gorilla algorithms.
  • Automated lifecycle retention: Time-series stores group data into discrete time chunks. Dropping data older than thirty days executes as a metadata directory drop rather than issuing expensive DELETE statements with index rebalancing.
  • Continuous downsampling: Engines aggregate raw one-second measurements into one-minute averages or hourly summaries in the background, conserving disk capacity while preserving historical trends.
Raw Metrics Stream (1-second intervals)
   |
   v (Fast append & gorilla compression)
[Hot Storage: 7-day retention]
   |
   v (Automated continuous rollups)
[Downsampled 1-hour aggregates: 1-year retention]

Balancing specialized engines against analytical warehouses is an essential design choice:

Retention cycles versus analytical warehouse queries

While time-series databases excel at serving sub-second queries for Grafana dashboards and alert evaluation rules, they are not general-purpose analytical warehouses. Time-series stores struggle when queries require relational joins across disparate business entities, complex nested subqueries, or cross-dataset exploratory scans.

Data platforms generally adopt a tiered pattern: telemetry streams flow directly into a time-series store like Prometheus or TimescaleDB to feed real-time monitoring and alerting. Batch pipelines or Kafka connectors export rolled-up daily metrics into columnar lakehouses like Snowflake or BigQuery for long-term historical analysis alongside customer and billing dimensions.

PreviousNext