Prometheus Metric Types & Instrumentation

Prometheus collects metrics: numeric time series identified by a name and labels. Applications expose them on an HTTP /metrics endpoint using a client library, and Prometheus scrapes them periodically. What you can later ask with PromQL depends entirely on how you instrumented: which metric types you chose, how you named them, and which labels you attached.

Good instrumentation is small, consistent, and cheap: a few well-chosen counters and histograms per service answer most operational questions. Bad instrumentation, with the wrong types, missing units, or labels holding user IDs, produces misleading graphs and cardinality explosions that can take down the monitoring system itself.

TL;DR

Quick Example

Instrumenting a Python web service with prometheus_client:

The resulting exposition format:

Core Concepts

Counter

A monotonically increasing value, reset to zero when the process restarts. Counters capture how many or how much happened. The raw value is useless; what matters is its rate:

Use counters for events: requests, errors, retries, bytes sent, cache hits and misses. Never decrement a counter, and don't use a gauge to count events: gauges miss increments between scrapes and don't handle restarts.

Gauge

A value that can go up or down, sampled at scrape time: memory usage, queue length, active sessions, pool connections in use, last successful run timestamp. Query gauges directly, or with avg_over_time, max_over_time, deriv, or predict_linear. Since gauges are sampled, short spikes between scrapes are invisible. Track peaks with a histogram or a max-tracking gauge if they matter.

Histogram

A histogram counts observations into configurable buckets (le = "less than or equal"), plus a running _sum and _count. Because buckets are counters, you can rate() and sum them across instances, then compute any quantile in PromQL with histogram_quantile. The trade-offs are that bucket boundaries must suit the value range (quantiles are interpolated within buckets) and each bucket is a separate series.

Native histograms (stable in Prometheus 3.x, supported by recent client libraries) use sparse exponential buckets stored as a single series. They offer much better resolution, no bucket tuning, and lower cost. Prefer them when your stack supports them.

Summary

A summary computes quantiles (for example p50 and p99) in the client over a sliding window and exposes them directly, along with _sum and _count. Summaries are accurate per instance, but quantiles can't be aggregated across instances (averaging p99s is meaningless), and the quantiles are fixed at instrumentation time. Use histograms unless you have a specific reason not to.

Info and State Metrics

Naming and Units

Labels and Cardinality

Each unique label combination is a separate time series with its own memory and storage cost:

That's fine. Add user_id with 1 million values and it's catastrophic.

Put high-cardinality details in logs and traces (see OpenTelemetry), and link them to metrics with exemplars.

What to Measure

Many metrics come free from exporters and frameworks. See Prometheus exporters.

Best Practices

Instrument at Boundaries

Measure where work enters and leaves your service: HTTP handlers, RPC clients, database calls, queue consumers. Middleware from web frameworks and OpenTelemetry instrumentation provides most of this automatically.

Initialize Label Combinations

Series only appear after their first observation, so an error counter that has never incremented doesn't exist, and rate() returns nothing rather than zero. Pre-initialize known label combinations to 0 where alerts depend on them.

Choose Buckets for Your SLOs

Include bucket boundaries at your latency objectives (for example 0.1s, 0.3s, 1s), so SLO queries are exact rather than interpolated. See SLOs. Or use native histograms.

Review Cardinality Regularly

Check the TSDB status page (/tsdb-status) or topk(10, count by (__name__) ({__name__=~".+"})) for metrics with exploding series counts, and enforce limits with sample_limit and relabeling in scrape configs.

Common Mistakes

Using Raw Paths as Labels

Milliseconds and Custom Units

request_latency_ms breaks conventions that tools rely on, and mixing units across services makes shared dashboards wrong. Use seconds as floats.

Summaries for Fleet-Wide Percentiles

Averaging per-instance p99s from summaries gives a number that isn't the fleet's p99. Use histograms and aggregate buckets.

FAQ

Counter or gauge?

If the value counts events that accumulate (requests, errors, bytes), use a counter and query its rate. If it's a current level that can decrease (queue length, memory, connections), use a gauge. A useful test: would rate() of it be meaningful? If yes, it's a counter.

Histogram or summary?

Histogram, in almost all cases: it's aggregatable across instances, you choose quantiles at query time, and it supports SLO calculations. Summaries only make sense for a single instance where you need very precise quantiles and can't design buckets, and native histograms remove even that reason.

How many labels is too many?

It's about the product of distinct values, not the count of labels. Keep total series per metric in the thousands to low hundreds of thousands, and never use labels whose values grow with users, requests, or time.

Should I use OpenTelemetry or Prometheus client libraries?

Both work. Prometheus client libraries are simple and mature for metrics-only needs. OpenTelemetry SDKs provide metrics, traces, and logs with one API and can export to Prometheus (scrape or OTLP remote write, which Prometheus 3.x accepts natively). New polyglot systems increasingly standardize on OpenTelemetry. See OpenTelemetry instrumentation.

Related Topics

References