A single IoT temperature sensor generates 86,400 data points per day at 1-second resolution. An application performance monitoring system ingesting metrics from 100 microservices produces millions of data points per minute. A financial trading platform stores tick data for thousands of instruments at millisecond granularity. These workloads share a characteristic that makes standard relational databases struggle: data arrives in continuous high-volume streams ordered by timestamp, queries filter and aggregate across time ranges, and historical data must be retained at progressively lower resolution as it ages.
Time series databases (TSDB) are purpose-built for this access pattern. Their write path is optimized for append-heavy, time-ordered inserts. Their query engine understands time-range filters natively. Their storage layer compresses time-correlated data 10–100x better than general-purpose formats. This guide covers the full TSDB landscape: what makes time-series data different, the four leading TSDB options, data modeling for time-series, downsampling and retention policies, and production deployment patterns for IoT and monitoring workloads.
Time Series Database: What Makes It Different
Time-series data has three characteristics that distinguish it from transactional data and motivate specialized storage:
Append-Heavy Write Pattern
Time-series data is almost exclusively append-only. New measurements arrive continuously; historical records are rarely updated or deleted. This property enables write-optimized storage structures — LSM trees (Log-Structured Merge trees) and columnar chunk storage — that batch writes for sequential disk I/O. A B-tree index (standard in relational databases) is optimized for random access patterns, not continuous sequential appends.
At high ingestion rates (100,000+ points/second), the difference between append-optimized and general-purpose write paths is orders of magnitude.
Time-Ordered Query Access
Queries almost always include a time range filter: "CPU usage for the last 24 hours," "temperature anomalies in March," "revenue trend for Q1 vs Q4." Time-based partitioning and time-ordered indexes make these range scans efficient. Without them, a query for "last 24 hours" on a 10-billion-row table performs a full scan.
Retention policies and downsampling — automatically summarizing and deleting old data — are standard TSDB capabilities. Historical data aged past 90 days might be downsampled from second-level to minute-level resolution, reducing storage by 60x while preserving trend visibility.
Aggregation-Heavy Queries
Time-series workloads compute aggregations: average, max, min, sum, percentiles, rates of change. These aggregations often run over millions or billions of data points. Columnar storage formats compress similar values efficiently and enable vectorized computation over columns rather than row-by-row processing.
Pre-aggregated "rollup" tables or continuous aggregates (materialized summaries at coarser time granularity) are the primary query performance technique — computing the average over 1-minute buckets is 60x faster than re-computing over raw second-level data.
TSDB Comparison: InfluxDB, TimescaleDB, Prometheus, QuestDB
InfluxDB
Purpose-built TSDB with its own storage engine (TSM — Time-Structured Merge tree), query language (Flux and InfluxQL), and ecosystem. InfluxDB 3.0 rewrites the storage engine using Apache Arrow and Parquet, significantly improving analytical query performance.
Data model: Measurements (like tables), tags (indexed metadata), fields (measured values), timestamp.
# InfluxDB line protocol
cpu_usage,host=server-01,region=us-east cpu_pct=87.3,mem_pct=62.1 1712000000000000000
temperature,sensor=sensor_42,location=building-a value=21.4 1712000000000000000
Strengths:
- Native TSDB with optimal write performance for pure time-series workloads
- Rich ecosystem: Telegraf (collection agent), Grafana integration, client libraries
- Built-in retention policies and continuous queries
- InfluxDB Cloud (managed, serverless pricing)
Limitations:
- Not relational — no JOINs with transactional data
- Flux query language has a steep learning curve
- Schema design mistakes are hard to correct post-ingestion
Ideal for: IoT sensor networks, application metrics, infrastructure monitoring, any use case that is 100% time-series with no relational requirements.
TimescaleDB
PostgreSQL extension that transforms PostgreSQL into a time-series database. Tables are called hypertables; TimescaleDB automatically partitions them by time (and optionally by a second dimension like device ID).
-- Create a hypertable (automatically partitioned by time)
CREATE TABLE sensor_readings (
time TIMESTAMPTZ NOT NULL,
sensor_id INTEGER NOT NULL,
temperature DOUBLE PRECISION,
humidity DOUBLE PRECISION
);
SELECT create_hypertable('sensor_readings', 'time',
chunk_time_interval => INTERVAL '1 day');
-- Optional: space partitioning for parallel ingestion
SELECT add_dimension('sensor_readings', 'sensor_id', number_partitions => 4);
-- Insert data (same as regular PostgreSQL)
INSERT INTO sensor_readings (time, sensor_id, temperature, humidity)
VALUES (NOW(), 42, 21.4, 58.2);
Continuous aggregates (materialized rollups):
-- Hourly average pre-computed continuously
CREATE MATERIALIZED VIEW sensor_hourly_avg
WITH (timescaledb.continuous) AS
SELECT
time_bucket('1 hour', time) AS hour,
sensor_id,
AVG(temperature) AS avg_temp,
AVG(humidity) AS avg_humidity,
MAX(temperature) AS max_temp,
MIN(temperature) AS min_temp
FROM sensor_readings
GROUP BY hour, sensor_id
WITH NO DATA;
-- Refresh policy
SELECT add_continuous_aggregate_policy('sensor_hourly_avg',
start_offset => INTERVAL '1 month',
end_offset => INTERVAL '1 hour',
schedule_interval => INTERVAL '1 hour');
Retention policies:
-- Automatically drop data older than 90 days
SELECT add_retention_policy('sensor_readings', INTERVAL '90 days');
Strengths:
- Full PostgreSQL compatibility — use all SQL features, JOINs, foreign keys
- Single database for time-series + relational business data
- Continuous aggregates for fast dashboard queries
- Managed on Timescale Cloud or self-hosted
Ideal for: Applications that need time-series AND relational queries in the same system; teams with PostgreSQL expertise; avoid dual-database complexity.
Prometheus
Pull-based metrics collection system with its own TSDB storage. The de facto standard for Kubernetes and infrastructure monitoring.
Data model: Metric name + label set + float64 value + timestamp.
# prometheus.yml — scrape targets
scrape_configs:
- job_name: 'application'
scrape_interval: 15s
static_configs:
- targets: ['app-01:9090', 'app-02:9090']
- job_name: 'kubernetes-pods'
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: true
PromQL queries:
# Request rate per service (5-minute window)
rate(http_requests_total{service="checkout"}[5m])
# 95th percentile latency
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
# Memory usage alert rule
ALERT HighMemoryUsage
IF process_resident_memory_bytes / node_memory_MemTotal_bytes * 100 > 90
FOR 5m
LABELS { severity = "warning" }
Limitations:
- 15-day default retention (Thanos or VictoriaMetrics for long-term storage)
- Pull-based model requires network access to scraped targets
- Not suitable for high-cardinality labels (millions of unique label combinations)
Ideal for: Kubernetes cluster monitoring, microservice metrics, alerting infrastructure (Alertmanager). Not suitable as a general IoT data store.
QuestDB
High-performance time series database optimized for analytical queries. Claims fastest ingestion speed in published benchmarks (1.6 million rows/second per CPU core). SQL-compatible with time-series extensions.
-- QuestDB SQL with time-series extensions
SELECT
timestamp,
sensor_id,
avg(temperature) OVER (
PARTITION BY sensor_id
ORDER BY timestamp
ROWS BETWEEN 5 PRECEDING AND CURRENT ROW
) AS rolling_5_avg
FROM sensor_readings
WHERE timestamp >= dateadd('d', -7, now())
SAMPLE BY 1h FILL(LINEAR);
-- SAMPLE BY: time-bucket aggregation
-- FILL(LINEAR): interpolate missing values
Ideal for: Analytical workloads over time-series data requiring SQL, financial tick data analysis, research and data science workflows.
TSDB Selection Decision Framework
| Use Case | Recommended Tool | Why |
|---|---|---|
| IoT sensor network, pure time-series | InfluxDB | Native TSDB, Telegraf ecosystem |
| Application + infrastructure monitoring | Prometheus + Grafana | Pull-based, Kubernetes native |
| Time-series + relational business data | TimescaleDB | PostgreSQL compatibility |
| High-performance analytical time-series SQL | QuestDB | Fastest analytical queries |
| Long-term Prometheus storage | VictoriaMetrics | Compatible, efficient compression |
Data Modeling for Time Series
Tag vs Field (InfluxDB) / Tag vs Column (TimescaleDB)
In InfluxDB: tags are indexed (use for filtering, group-by); fields are not indexed (use for measured values). Tags are metadata you filter on; fields are values you measure.
# Good: location is a tag (filter by it), temperature is a field (measure it)
temperature,sensor=42,location=building-a value=21.4
# Bad: using a field as a filter dimension
temperature sensor=42,location=building-a,temperature=21.4
High-cardinality tag problem: Tags with millions of distinct values (unique request IDs, individual user IDs) create millions of series, degrading InfluxDB performance. In Prometheus, high-cardinality labels cause memory explosions. Design tags with bounded cardinality: device type, region, service name — not individual entity IDs.
Downsampling Strategy
Raw second-level data is expensive to store and slow to query for long time ranges. Downsampling converts raw data to aggregates at coarser granularity:
| Age | Resolution | Storage Reduction |
|---|---|---|
| 0–7 days | 1-second (raw) | Baseline |
| 7–90 days | 1-minute averages | 60x |
| 90 days – 2 years | 1-hour averages | 3,600x |
| 2+ years | 1-day averages | 86,400x |
This tiered retention strategy enables answering "what happened in the last hour?" with high precision while still enabling "what was the trend last year?" without petabytes of raw storage.
InfluxDB vs TimescaleDB: Detailed Comparison
The most common TSDB selection decision is between InfluxDB and TimescaleDB. The right choice depends on whether you need SQL compatibility and relational joins.
Query language: InfluxDB uses Flux (InfluxDB 2.x/3.x) and InfluxQL (legacy). Flux is a functional data scripting language specifically designed for time-series transformations — powerful for time-series-specific operations but requires learning new syntax. TimescaleDB uses standard SQL with time-series extensions (time_bucket, first, last, locf). If your team knows SQL, TimescaleDB's learning curve is near-zero.
Joins with relational data: InfluxDB has no JOIN capability — measurements exist in isolation. TimescaleDB can JOIN hypertables with regular PostgreSQL tables:
-- TimescaleDB: join sensor readings with device metadata
SELECT
d.device_name,
d.location,
time_bucket('1 hour', sr.time) AS hour,
AVG(sr.temperature) AS avg_temp
FROM sensor_readings sr
JOIN devices d ON sr.sensor_id = d.id
WHERE sr.time >= NOW() - INTERVAL '7 days'
AND d.device_type = 'temperature_sensor'
GROUP BY d.device_name, d.location, hour
ORDER BY hour DESC;
This is impossible in InfluxDB without pre-joining data at ingestion time.
Schema flexibility: InfluxDB is schemaless — new fields can be added to any measurement without schema changes. TimescaleDB requires column definitions, like any PostgreSQL table. For IoT workloads where sensor types and reported fields change frequently, InfluxDB's schemaless model reduces friction. For monitoring with a stable metric schema, the difference is minor.
Ingestion performance: InfluxDB's native write protocol (line protocol) is optimized for high-frequency time-series ingestion. TimescaleDB uses PostgreSQL's write path with chunk-based optimization — excellent performance for most use cases, but InfluxDB has an edge for sustained extremely high ingestion rates (millions of points per second).
Operational complexity: TimescaleDB runs as a PostgreSQL extension. If you already run PostgreSQL, TimescaleDB is an apt install and CREATE EXTENSION timescaledb away. InfluxDB is a separate service requiring separate deployment, backup, and monitoring.
Production Deployment Patterns
Write path: Buffer high-frequency writes in an in-memory queue (or Kafka topic) before bulk-flushing to the TSDB. Most TSDBs accept batch inserts efficiently; single-row inserts at high rates introduce unnecessary per-request overhead.
Grafana integration: Grafana connects natively to InfluxDB, TimescaleDB, Prometheus, and QuestDB. Configure variable-based dashboard templates that allow users to select time range, resolution, and filter dimensions (device, region, service) without custom code.
Capacity planning: Time-series data volume is predictable. sensors × data_points_per_second × bytes_per_point × retention_days. A network of 10,000 sensors at 1 reading/second for 90 days at 50 bytes/reading = 3.9TB of raw data. Apply downsampling and compression ratios (typically 10–20x for TSDB) to get actual storage requirements.
Conclusion
Time series databases solve a real problem: general-purpose databases struggle with high-frequency append-only data, time-range queries, and long retention with downsampling. The right TSDB for your use case depends on whether you need relational joins (TimescaleDB), a monitoring ecosystem (Prometheus), a managed pure-TSDB (InfluxDB Cloud), or analytical SQL (QuestDB).
For most teams adding time-series capability to an existing PostgreSQL system, TimescaleDB is the lowest-friction option — it adds time-series optimization without introducing a new operational system to manage. For pure monitoring infrastructure or IoT networks with no relational requirements, InfluxDB's native TSDB optimizations and ecosystem provide the best fit.
Author: Smart Maple Database Engineering Team Updated: April 2026
Related Articles
MLOps Guide: Taking Machine Learning Models to Production [2026]
87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow
Read MoreLLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]
General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not
Read MoreComputer Vision Applications: Object Detection, OCR, and Industrial AI [2026]
Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl
Read More
