Skip to content

Prometheus

On this page

The metrics and alerting workhorse: the pull model, the time-series data model, PromQL, exporters, recording rules, and Alertmanager.

Basics

Prometheus (CNCF) is an open-source metrics-based monitoring and alerting system. It scrapes (pulls) metrics over HTTP from targets at intervals, stores them as time series in a local TSDB, and lets you query with PromQL.

The data model

A time series is identified by a metric name + a set of labels (key/value):

http_requests_total{method="GET", handler="/api", status="200"}  →  value @ timestamp

Labels make metrics multi-dimensional; every unique label combination is a separate series.

Metric types

Type Use Example
Counter Monotonically increasing total http_requests_total
Gauge Goes up and down node_memory_MemAvailable_bytes
Histogram Bucketed observations + sum/count request_duration_seconds_bucket
Summary Client-side quantiles rpc_duration_seconds

Architecture

 targets (apps + exporters)  ──/metrics──▶  Prometheus server
                                              │  scrape + store (TSDB)
                                              │  evaluate rules
                          PromQL ◀── Grafana  │
                                              ▼
                                        Alertmanager ──▶ Slack/Email/PagerDuty
  • Exporters translate third-party systems into Prometheus metrics (node_exporter for hosts, blackbox_exporter for probes, plus DB/queue/app exporters).
  • Pushgateway — for short-lived batch jobs that can't be scraped (use sparingly).
  • Service discovery — auto-find targets in Kubernetes, EC2, Consul, etc.
  • Alertmanager — dedupes, groups, silences, and routes alerts to receivers.
  • Remote write — ship samples to long-term stores (Thanos, Mimir, Cortex).

Pull vs push

Prometheus pulls by default: it controls scrape timing, can detect a target being down (up == 0), and avoids clients overwhelming the server. Push is the exception (batch jobs).

Cheatsheet

Minimal prometheus.yml

global:
  scrape_interval: 15s
  evaluation_interval: 15s

rule_files:
  - rules/*.yml

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["alertmanager:9093"]

scrape_configs:
  - job_name: prometheus
    static_configs:
      - targets: ["localhost:9090"]

  - job_name: node
    static_configs:
      - targets: ["10.0.0.11:9100", "10.0.0.12:9100"]

  - job_name: kubernetes-pods       # example SD
    kubernetes_sd_configs:
      - role: pod

PromQL essentials

# Per-second request rate over 5m (counters → always use rate)
rate(http_requests_total[5m])

# Total across labels
sum(rate(http_requests_total[5m])) by (status)

# Error ratio
sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))

# CPU usage per instance (from node_exporter)
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)

# Memory available %
node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100

# 95th percentile latency from a histogram
histogram_quantile(0.95, sum(rate(request_duration_seconds_bucket[5m])) by (le))

# Is the target up?
up == 0

# Predict disk full within 4h
predict_linear(node_filesystem_avail_bytes[6h], 4*3600) < 0

Recording & alerting rules

groups:
  - name: example
    rules:
      - record: job:http_requests:rate5m          # precompute heavy queries
        expr: sum(rate(http_requests_total[5m])) by (job)

      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)
            / sum(rate(http_requests_total[5m])) by (job) > 0.05
        for: 10m
        labels: { severity: critical }
        annotations:
          summary: "High 5xx error rate on {{ $labels.job }}"

Useful endpoints

http://prometheus:9090/             # expression browser
/targets                            # scrape health
/rules  /alerts                     # rules & firing alerts
promtool check config prometheus.yml
promtool check rules rules/*.yml

Thumb Rules

Rules of thumb

  • rate()/increase() on counters, never raw counters. Counters reset on restart.
  • Watch label cardinality. High-cardinality labels (user IDs, request IDs) explode series and memory — keep labels bounded.
  • Use for: on alerts so transient blips don't page you.
  • Histograms for latency, then histogram_quantile — don't average percentiles.
  • Record expensive queries so dashboards/alerts stay fast.
  • Prometheus storage is local & not for long-term/HA — use remote-write (Thanos/Mimir) for that.
  • Alert on symptoms, not causes (user-facing SLOs) — see Google's "four golden signals": latency, traffic, errors, saturation.
  • up == 0 is your most important alert — a target you can't scrape is invisible.

Use Cases

  • Infrastructure monitoring — hosts (node_exporter), containers (cAdvisor), Kubernetes.
  • Application/RED metrics — Rate, Errors, Duration via client libraries.
  • SLO/SLA tracking — error budgets from histograms and ratios.
  • Capacity planning — trends and predict_linear.
  • Blackbox probing — HTTP/TCP/DNS/ICMP endpoint checks.
  • Alerting backbone — feeds Alertmanager → on-call.

Common Issues

Target shows DOWN on /targets

Network/firewall to the /metrics port, wrong scheme, or the exporter isn't running. Curl the target's /metrics from the Prometheus host; check the error in /targets.

High memory usage / OOM

Cardinality explosion — too many label combinations. Find the worst metrics with topk(20, count by (__name__)({__name__=~".+"})), then drop or relabel high-cardinality labels (user IDs, request IDs, emails) at scrape time.

Counter looks wrong / goes down

The process restarted (counters reset to 0). Always wrap in rate()/increase(), which handle resets.

Alerts not firing / not delivered

Check /alerts (is it pending/firing?), the for: duration, and Alertmanager routing. Validate with promtool check rules and inspect Alertmanager logs.

Slow dashboards / queries time out

Heavy PromQL over long ranges with high cardinality. Add recording rules, shorten ranges, and reduce series.

Data gaps after restart

Single-node Prometheus has downtime during restarts and limited retention. Use HA pairs and remote-write for durability.

Best Practices

  • Standardize metric & label naming (_total, _seconds, base units) per the official guidelines.
  • Keep cardinality under control; relabel/drop noisy labels at scrape time.
  • Instrument the RED/USE methods; align alerts to SLOs and the four golden signals.
  • Use recording rules for dashboards and frequent alert expressions.
  • Run Alertmanager with sensible grouping, inhibition, and silences; route by severity.
  • Plan retention & long-term storage (Thanos/Mimir/Cortex) and HA scrapers.
  • Secure endpoints (auth/TLS/reverse proxy); /metrics can leak internals.
  • Validate config in CI with promtool.

Official Sources