Prometheus¶
On this page
The metrics and alerting workhorse: the pull model, the time-series data model, PromQL, exporters, recording rules, and Alertmanager.
Basics¶
Prometheus (CNCF) is an open-source metrics-based monitoring and alerting system. It scrapes (pulls) metrics over HTTP from targets at intervals, stores them as time series in a local TSDB, and lets you query with PromQL.
The data model¶
A time series is identified by a metric name + a set of labels (key/value):
Labels make metrics multi-dimensional; every unique label combination is a separate series.
Metric types¶
| Type | Use | Example |
|---|---|---|
| Counter | Monotonically increasing total | http_requests_total |
| Gauge | Goes up and down | node_memory_MemAvailable_bytes |
| Histogram | Bucketed observations + sum/count | request_duration_seconds_bucket |
| Summary | Client-side quantiles | rpc_duration_seconds |
Architecture¶
targets (apps + exporters) ──/metrics──▶ Prometheus server
│ scrape + store (TSDB)
│ evaluate rules
PromQL ◀── Grafana │
▼
Alertmanager ──▶ Slack/Email/PagerDuty
- Exporters translate third-party systems into Prometheus metrics
(
node_exporterfor hosts,blackbox_exporterfor probes, plus DB/queue/app exporters). - Pushgateway — for short-lived batch jobs that can't be scraped (use sparingly).
- Service discovery — auto-find targets in Kubernetes, EC2, Consul, etc.
- Alertmanager — dedupes, groups, silences, and routes alerts to receivers.
- Remote write — ship samples to long-term stores (Thanos, Mimir, Cortex).
Pull vs push¶
Prometheus pulls by default: it controls scrape timing, can detect a target being down
(up == 0), and avoids clients overwhelming the server. Push is the exception (batch jobs).
Cheatsheet¶
Minimal prometheus.yml¶
global:
scrape_interval: 15s
evaluation_interval: 15s
rule_files:
- rules/*.yml
alerting:
alertmanagers:
- static_configs:
- targets: ["alertmanager:9093"]
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ["localhost:9090"]
- job_name: node
static_configs:
- targets: ["10.0.0.11:9100", "10.0.0.12:9100"]
- job_name: kubernetes-pods # example SD
kubernetes_sd_configs:
- role: pod
PromQL essentials¶
# Per-second request rate over 5m (counters → always use rate)
rate(http_requests_total[5m])
# Total across labels
sum(rate(http_requests_total[5m])) by (status)
# Error ratio
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# CPU usage per instance (from node_exporter)
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
# Memory available %
node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100
# 95th percentile latency from a histogram
histogram_quantile(0.95, sum(rate(request_duration_seconds_bucket[5m])) by (le))
# Is the target up?
up == 0
# Predict disk full within 4h
predict_linear(node_filesystem_avail_bytes[6h], 4*3600) < 0
Recording & alerting rules¶
groups:
- name: example
rules:
- record: job:http_requests:rate5m # precompute heavy queries
expr: sum(rate(http_requests_total[5m])) by (job)
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)
/ sum(rate(http_requests_total[5m])) by (job) > 0.05
for: 10m
labels: { severity: critical }
annotations:
summary: "High 5xx error rate on {{ $labels.job }}"
Useful endpoints¶
http://prometheus:9090/ # expression browser
/targets # scrape health
/rules /alerts # rules & firing alerts
promtool check config prometheus.yml
promtool check rules rules/*.yml
Thumb Rules¶
Rules of thumb
rate()/increase()on counters, never raw counters. Counters reset on restart.- Watch label cardinality. High-cardinality labels (user IDs, request IDs) explode series and memory — keep labels bounded.
- Use
for:on alerts so transient blips don't page you. - Histograms for latency, then
histogram_quantile— don't average percentiles. - Record expensive queries so dashboards/alerts stay fast.
- Prometheus storage is local & not for long-term/HA — use remote-write (Thanos/Mimir) for that.
- Alert on symptoms, not causes (user-facing SLOs) — see Google's "four golden signals": latency, traffic, errors, saturation.
up == 0is your most important alert — a target you can't scrape is invisible.
Use Cases¶
- Infrastructure monitoring — hosts (node_exporter), containers (cAdvisor), Kubernetes.
- Application/RED metrics — Rate, Errors, Duration via client libraries.
- SLO/SLA tracking — error budgets from histograms and ratios.
- Capacity planning — trends and
predict_linear. - Blackbox probing — HTTP/TCP/DNS/ICMP endpoint checks.
- Alerting backbone — feeds Alertmanager → on-call.
Common Issues¶
Target shows DOWN on /targets
Network/firewall to the /metrics port, wrong scheme, or the exporter isn't running.
Curl the target's /metrics from the Prometheus host; check the error in /targets.
High memory usage / OOM
Cardinality explosion — too many label combinations. Find the worst metrics with
topk(20, count by (__name__)({__name__=~".+"})), then drop or relabel
high-cardinality labels (user IDs, request IDs, emails) at scrape time.
Counter looks wrong / goes down
The process restarted (counters reset to 0). Always wrap in rate()/increase(), which
handle resets.
Alerts not firing / not delivered
Check /alerts (is it pending/firing?), the for: duration, and Alertmanager routing.
Validate with promtool check rules and inspect Alertmanager logs.
Slow dashboards / queries time out
Heavy PromQL over long ranges with high cardinality. Add recording rules, shorten ranges, and reduce series.
Data gaps after restart
Single-node Prometheus has downtime during restarts and limited retention. Use HA pairs and remote-write for durability.
Best Practices¶
- Standardize metric & label naming (
_total,_seconds, base units) per the official guidelines. - Keep cardinality under control; relabel/drop noisy labels at scrape time.
- Instrument the RED/USE methods; align alerts to SLOs and the four golden signals.
- Use recording rules for dashboards and frequent alert expressions.
- Run Alertmanager with sensible grouping, inhibition, and silences; route by severity.
- Plan retention & long-term storage (Thanos/Mimir/Cortex) and HA scrapers.
- Secure endpoints (auth/TLS/reverse proxy);
/metricscan leak internals. - Validate config in CI with
promtool.
Official Sources¶
- Prometheus Documentation — https://prometheus.io/docs/introduction/overview/
- Querying / PromQL — https://prometheus.io/docs/prometheus/latest/querying/basics/
- Metric & label naming — https://prometheus.io/docs/practices/naming/
- Alerting & Alertmanager — https://prometheus.io/docs/alerting/latest/overview/
- Exporters list — https://prometheus.io/docs/instrumenting/exporters/
- Node Exporter — https://github.com/prometheus/node_exporter
- Google SRE Book (golden signals) — https://sre.google/sre-book/monitoring-distributed-systems/
- Thanos (long-term/HA) — https://thanos.io/