Prometheus Alerting & Alertmanager
Prometheus alerting has two parts. Alerting rules, evaluated by Prometheus, are PromQL expressions that fire when a condition holds for long enough. Alertmanager receives those firing alerts and decides what to do with them: it groups related alerts, deduplicates them across Prometheus replicas, routes them to the right team and channel, inhibits noisy downstream alerts, and supports silences during maintenance.
The technology is straightforward. The hard part is alert design: paging only on problems users feel, with enough context to act, and without waking people for noise. Good alerting follows the practices of SRE: alert on symptoms and SLO burn, not every cause.
TL;DR
- Alerting rules are PromQL expressions plus
for(how long before firing), labels (severity, team), and annotations (summary, runbook). - Alert on symptoms users experience (errors, latency, availability), and use dashboards for causes.
- Multi-window burn-rate alerts on SLOs catch both fast outages and slow degradation with few false positives.
- Alertmanager routes by labels, groups related alerts, inhibits dependent ones, and applies silences.
- Every page needs a runbook and must be actionable; everything else is a ticket or a dashboard.
- Run Alertmanager as an HA cluster and alert on the alerting pipeline itself (a watchdog or dead man's switch).
Quick Example
Alerting rules (loaded by Prometheus):
Alertmanager routing:
Core Concepts
Alerting Rules
Prometheus evaluates each rule group every evaluation_interval. An alert is:
- Inactive: the expression returns nothing.
- Pending: the expression returns series, but not yet for the full
forduration. - Firing: it has been true for at least
for; Prometheus sends it to Alertmanager.
for filters out brief blips. keep_firing_for (Prometheus 2.42+) keeps an alert firing for a while after the condition clears, which prevents flapping. Labels drive routing (severity, team, service). Annotations give responders context, templated with {{ $labels.x }} and {{ $value }}.
Symptom-Based Alerting
Page on what users feel: elevated error rate, high latency, unavailability, and data freshness for pipelines. Causes (high CPU, a pod restart, disk at 70%) belong on dashboards or in low-priority tickets, unless they predict imminent user impact (disk full within hours via predict_linear).
SLO Burn-Rate Alerts
A burn rate is how fast you're consuming your error budget relative to plan. With a 99.9% SLO, the budget is 0.1% errors, and a burn rate of 14.4 exhausts a 30-day budget in about 2 days. The Google SRE multi-window, multi-burn-rate approach:
The short window makes alerts resolve quickly once the problem stops. Tools like Sloth and Pyrra generate these rules from SLO definitions.
Alertmanager Concepts
- Routing tree: alerts match routes by label matchers. Children inherit settings, and
continue: truelets an alert match several routes. - Grouping:
group_bycombines alerts with the same labels into one notification.group_waitdelays the first notification to collect related alerts,group_intervalcontrols updates, andrepeat_intervalcontrols reminders for still-firing alerts. - Inhibition: suppress target alerts while a source alert fires with matching
equallabels. For example, no per-service pages when a whole cluster or datacenter is down. - Silences: time-bounded mutes by matcher, created in the UI or with
amtool, for maintenance and known issues. - Receivers: Slack, PagerDuty, Opsgenie, email, Microsoft Teams, webhooks, and more, with templated messages.
- High availability: run several Alertmanagers in a gossip cluster. Every Prometheus sends to all of them, and they deduplicate notifications.
Best Practices
Make Every Page Actionable
If the responder can't do anything about it right now, it shouldn't page. Downgrade to a ticket or a dashboard, or delete it. Review pages after every on-call shift, and prune noisy alerts relentlessly. See incident management.
Include a Runbook and a Dashboard
Annotations should link to a runbook (what it means, how to diagnose, how to mitigate) and a dashboard scoped to the problem. That context turns 20 minutes of orientation into 2.
Test Alert Rules
Validate syntax with promtool check rules, and write unit tests with promtool test rules, feeding synthetic series and asserting which alerts fire when. Keep rules in version control and deploy via CI.
Monitor the Monitoring
Run an always-firing Watchdog alert routed to a dead man's switch service (such as healthchecks.io or PagerDuty's heartbeat), so you're notified if Prometheus or Alertmanager stops working. Alert on failed notifications (alertmanager_notifications_failed_total) and rule evaluation failures.
Common Mistakes
Alerting on Every Cause
Page on latency and errors instead, and investigate CPU from the dashboard when they fire.
No for Duration
Without for, a single bad scrape or momentary spike pages someone. Use for (or burn-rate windows) proportional to how quickly you need to react.
Alerts That Vanish When Data Disappears
rate(errors_total[5m]) > 1 can't fire if the service is down and not exporting metrics at all. Pair symptom alerts with up == 0 or absent() alerts for critical jobs.
FAQ
What's the difference between Prometheus alerting rules and Alertmanager?
Prometheus evaluates rules and decides whether an alert is firing. Alertmanager decides who gets notified, how, and when: grouping, routing, deduplication, inhibition, and silencing. Separating them lets multiple Prometheus servers share one notification pipeline.
How do I avoid duplicate notifications from HA Prometheus pairs?
Configure both Prometheus replicas to send alerts to the same Alertmanager cluster. Identical alerts from both are deduplicated automatically. Make sure external labels like replica are dropped (via alert_relabel_configs) so the alerts look identical.
Should I alert in Grafana or Alertmanager?
Both work. Grafana Alerting can query many data sources and has a friendly UI; Prometheus rules plus Alertmanager are file-based, version-controlled, and evaluated close to the data. Many teams keep critical SLO alerts as Prometheus rules in Git, and use Grafana alerts for cross-source or ad hoc cases.
What severity levels should we use?
Keep it simple: page (wake someone up: user impact now or imminent), ticket (needs attention within business hours), and optionally info (dashboards and chat only). More levels usually add confusion rather than clarity.
Related Topics
- Prometheus — The monitoring system overview
- PromQL — Writing alert expressions
- Alerting — Alert design principles
- SLOs — Error budgets and burn rates
- Incident Management — What happens after the page
- Runbooks — Giving responders a starting point