Prometheus

Prometheus is an open-source monitoring system that collects numeric time-series metrics by periodically scraping HTTP endpoints, stores them in a local time-series database, and lets you query them with PromQL. It graduated as the second CNCF project after Kubernetes and is the de facto standard for metrics in cloud-native systems.

Prometheus does one job — metrics — very well. It pairs with Grafana for dashboards, Alertmanager for notifications, and tools like Thanos or Mimir for long-term, multi-cluster storage. Logs and traces are handled by other tools in the observability stack.

TL;DR

Quick Example

Instrument a Node.js service with prom-client, exposing a request counter and a latency histogram:

Then ask Prometheus for the error ratio and p95 latency:

Core Concepts

The Pull Model and Service Discovery

Prometheus scrapes each target's /metrics endpoint every scrape interval (commonly 15–30 seconds). Targets come from static config or service discovery — Kubernetes, EC2, Consul, DNS. Pulling makes target health visible (a failed scrape sets up == 0) and keeps control over load on the Prometheus side. For short-lived batch jobs that end before a scrape, use the Pushgateway sparingly.

On Kubernetes, the Prometheus Operator (often via the kube-prometheus-stack Helm chart) manages this through ServiceMonitor and PodMonitor resources.

Metric Types

Histograms can be aggregated across pods; summaries can't — which is why histograms are the default for latency. Newer Prometheus versions also support native histograms with automatic, high-resolution buckets.

Labels and Cardinality

Each unique set of label values creates a new time series. A metric with route (20 values) × status (5) × pod (30) is 3,000 series — fine. Add user_id with a million values and you have billions, which exhausts memory. Cardinality is the main scaling limit of Prometheus.

PromQL Essentials

Rules and Alertmanager

Recording rules precompute expensive queries into new series. Alerting rules evaluate a PromQL condition and fire when it holds for a duration:

Alertmanager deduplicates, groups related alerts, applies silences and inhibitions, and routes notifications to PagerDuty, Slack, or email. See Alerting & On-Call.

Storage and Scaling

A single Prometheus server stores data locally, typically for 15–30 days, and scales vertically to millions of active series. For long retention, global queries across clusters, and high availability, add Thanos, Grafana Mimir, Cortex, or VictoriaMetrics, which receive data via remote_write or sidecars and store it in object storage.

Best Practices

Follow Naming Conventions

Use base units and suffixes: _seconds, _bytes, _total for counters. Consistent names make queries and dashboards predictable.

Instrument the Golden Signals First

Latency, traffic, errors, and saturation for every service catch most problems. The RED method (rate, errors, duration) for services and USE method (utilization, saturation, errors) for resources are good checklists.

Use Route Templates, Not Raw Paths

Label /orders/:id, not /orders/81723. Raw paths explode cardinality.

Alert on Symptoms and SLOs

Page on user-facing error rates and latency burning your SLO budget, not on CPU spikes. Every page should link a runbook.

Precompute Dashboard Queries

Recording rules keep dashboards fast and alert expressions readable.

Monitor Prometheus Itself

Watch prometheus_tsdb_head_series, scrape durations, and rule evaluation failures so you notice cardinality growth before it becomes an outage.

Common Mistakes

High-Cardinality Labels

Put per-request detail in logs or traces.

Graphing Raw Counters

A raw counter just climbs forever and resets on restart. Wrap it in rate() or increase().

Averaging Percentiles

Averaging p95 values across pods is mathematically wrong. Aggregate histogram buckets first, then apply histogram_quantile().

rate() Windows Shorter Than Two Scrapes

With a 30-second scrape interval, rate(x[30s]) often has too few samples. Use at least four times the scrape interval.

Using Prometheus as an Event Store

Prometheus is for aggregated numeric metrics, not per-event records or billing-grade accuracy.

FAQ

What is Prometheus used for?

Prometheus collects and stores numeric metrics from applications and infrastructure, lets you query them with PromQL, powers dashboards, and evaluates alerting rules. It's the standard metrics system for Kubernetes and cloud-native services.

Why does Prometheus pull instead of push?

Pulling lets Prometheus control scrape load, detect down targets automatically through failed scrapes, and discover targets dynamically. Push is still available through the Pushgateway or remote_write for special cases like short-lived batch jobs.

What is the difference between Prometheus and Grafana?

Prometheus collects, stores, and queries metrics and evaluates alerts. Grafana visualizes data from Prometheus and many other sources in dashboards. They're usually used together.

How long does Prometheus keep data?

The default local retention is 15 days, configurable by time or size. For months or years of history, send data to a long-term backend like Thanos, Mimir, or a managed Prometheus service.

Does Prometheus handle logs and traces?

No. Prometheus is metrics-only. Use Loki or Elasticsearch for logs and Tempo or Jaeger for traces, often instrumented through OpenTelemetry.

Related Topics

References